Research on steering behavior in 35B MoE language models using sparse autoencoders and probe vectors to identify and control agentic traits.
MLP architecture with learned structural dropout and input-dependent gating for conditional computation and regularization.
Federated learning framework for non-IID distributed scenarios using generative one-shot learning without foundation model dependencies.
Spectral initialization method for neural networks designed for function parameterization using prior information.
Methods for adding persistent memory to frozen encoder-decoder LLMs using continuous latent space adapters for multi-session learning.
Solver for distributional counterfactual explanations using optimal transport with statistical certification for model interpretability.
LLM compression method using capability-guided budget allocation that interprets what model components encode before pruning.
High-frequency time series dataset at millisecond resolution for training and evaluating time series foundation models.
Test-time scaling and confidence calibration strategy using internal model information for improved reinforcement learning.
Foundation model for structured data with linear complexity for handling extremely large datasets in healthcare, finance, and e-commerce.
Unsupervised autoencoder regularization by aligning pairwise distances between latent and input spaces on learned manifolds.
Deep learning methods for tabular data using representation correction to improve on in-learning and pre-learning paradigms.
Analysis of when unsupervised reinforcement learning succeeds in LLM mathematical reasoning, addressing scalability of outcome-based RL.
Method for discrete reasoning using self-aware Markov models that correct errors in masked diffusion models through adaptive denoising.
Study on how Transformers develop internal geometric representations of grid-world environments through next-token prediction.
Study showing chain-of-thought prompting degrades uncertainty quantification in vision-language models despite improving reasoning.
Analysis of quantized optimizer states in LLM pre-training, studying state staleness and effectiveness of reset strategies.
SpecMoE mixture-of-experts foundation model for cross-species EEG decoding with spectral and temporal signal analysis.
Contextual bandit algorithm combining dense arm features, non-linear rewards, and time-varying correlation for recommendations.
pADAM generative framework learning shared probabilistic priors across heterogeneous PDE families for multi-physics simulation.
SOMP algorithm for scaling gradient inversion attacks on LLMs, revealing privacy risks from shared gradients in large batch settings.
Conservative stochastic control framework for treatment optimization from irregularly sampled medical patient trajectories.
Method using adaptive moment estimation to stabilize guided diffusion sampling for image restoration and generation tasks.
Research on Gaussian mean estimation under realizable contamination with missing data patterns.
Stochastic resetting mechanism accelerates policy convergence in reinforcement learning on tabular environments.
Dynamic meta-layer aggregation defends federated learning against Byzantine adversaries and untargeted attacks.
Efficient chain-of-thought reasoning for edge deployment via compressed reasoning traces and smaller model distillation.
NextMem introduces latent factual memory for LLM-based agents, addressing catastrophic forgetting and context overhead.
Self-reflective recursive program search improves long-context handling in language models through programmatic decomposition.
Spiking neural networks for mobile robotics; biologically-inspired learning for power-constrained environments.
MiroThinker-1.7 and H1 research agents with verification for complex long-horizon reasoning and multi-step problem solving.
ClawWorm: self-propagating attack demonstrating security vulnerabilities in multi-agent LLM ecosystems like OpenClaw.
Simulation Distillation approach for sim-to-real transfer in robotics; pretrains world models for rapid real-world adaptation.
Theoretical characterization of partial labels learning feasibility with adaptive nearest neighbors method.
Behavioral Foundation Models baseline using regularized latent dynamics prediction for adaptive agent policies.
Theoretical analysis of transformers for knowledge retrieval in LLMs beyond orthogonal embedding assumptions.
Research on dynamic tokenization replacing fixed vocabularies in LLMs; hierarchical autoregressive approach for 70B parameter models.
Constitutional AI research on learning natural language rules automatically for LLM control via multi-agent framework.
Data augmentation framework using pseudo-labeling and unlabeled speech for robust dysarthric speech severity assessment.
Asymmetric pruning technique for vision-language models addressing modality-specific behaviors in text and visual token compression.
LLM-based framework using retrieval augmentation and confidence-based automation for efficient radiology report annotation in clinical NLP.
Power analysis framework for statistical inference on ML-predicted outcomes, addressing sample size determination for prediction-powered inference.
Feature selection method for distributionally robust learning maintaining reliability across diverse deployment environments with covariate shift.
Attribution upsampling method using redistribution instead of interpolation to prevent corruption of saliency maps in explainable AI.
Parallel in-context learning technique for vision-language models reducing inference latency while maintaining demonstration effectiveness.
Study showing LLM pre-training without learning rate decay improves downstream supervised fine-tuning performance.
Benchmark comparing GAN and Stable Diffusion augmentation strategies for class imbalance correction in animal classification under low-data conditions.
Graph-based multi-agent reinforcement learning for decentralized UAV swarm coordination under partial observability and communication constraints.
Deep Adaptive Design for efficient model-based design of experiments in nonlinear dynamical systems with offline neural network policies.
LLM-based recommender system using review aggregation and multi-factor attention for restaurant recommendations.