Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
LLM-based agent (GOME) for machine learning engineering using gradient-based reasoning instead of tree search for efficient optimization.
LLM-based agent (GOME) for machine learning engineering using gradient-based reasoning instead of tree search for efficient optimization.
Mechanistic interpretability study using architectural topology modifications to investigate grokking and delayed generalization in Transformers.
Study on eliciting truthful responses from censored LLMs and detecting deception using naturally-occurring dishonesty patterns.
Memory-efficient optimization method for training large language models with mask traversal and convergence guarantees for nonconvex settings.
Clustering algorithm for data summarization using Khatri-Rao product to reduce redundancies in centroid-based cluster prototypes.
Meta-optimizer that dynamically selects update rules during training with PyTorch compatibility, achieving 5.3x faster convergence.
Multi-objective protein sequence design balancing designability with competing properties like solubility and thermostability using preference alignment.
Reinforcement learning approach for robust policies under latent distribution shift in partially observable domains using adversarial training.
Generative models with tunable complexity for inverse problems, extending diffusion models and normalizing flows with adaptive dimensionality.
Federated learning framework addressing non-IID data distribution through personalization strategies for heterogeneous client data.
Benchmark for evaluating time-series forecasting foundation models with temporal generalization, addressing train-test contamination issues in static splits.
Deep learning method for learning coordinates and flow maps to enhance computational efficiency in multiscale dynamical systems.
DRUPI: dataset reduction method using privileged information beyond input-label pairs for dataset condensation tasks.
Differentiable optimization approach using control barrier functions to learn safe responsibility allocations in multi-agent autonomous systems.
Adaptive importance sampling and stratified subsampling estimators for robust high-dimensional sparse regression with heavy-tailed noise.
Prognostics system for autonomous deep-space habitat health monitoring and remaining useful life prediction under multiple failure modes.
MS-HGNN: morphological-symmetry-equivariant heterogeneous graph neural network for robotic dynamics learning with structural priors.
CuriousBot: mobile robotic exploration system using actionable 3D relational object graphs for interactive environment exploration.
Real2Sim2Real framework using likelihood-free inference for visual robotic manipulation of deformable linear objects.
LayerNorm tuning method using concept drift for efficient multimodal metaphor identification in internet memes.
UltraEdit: training-free model editing approach for lifelong knowledge updates in LLMs without retraining or subject-specific data.
CORA: cooperative game-theoretic credit assignment method for multi-agent reinforcement learning using core concepts for advantage allocation.
Regret-optimal Q-learning algorithms minimizing sample collection and policy switching costs in single-agent and federated reinforcement learning.
Supervised contrastive learning approach for low-resource language identification to improve multilingual LLM pretraining corpus curation.
Research on convergence rates for stochastic gradient descent and heavy ball methods under convex and non-convex objectives using discrete Gronwall's inequality.
Latent policy steering approach using embodiment-agnostic pretrained world models to leverage cross-embodiment datasets for robot learning.
Robot Control Stack ecosystem for scalable vision-language-action model training and deployment, replacing traditional robotics frameworks.
Test-time composition method for enhancing diffusion-based robot control policies without additional training data or model fine-tuning.
Latent Speech-Text Transformer improving compute efficiency of speech-text models by reducing token sequence length through latent representations.
AlphaApollo agentic reasoning system addressing reasoning capacity and verification bottlenecks through multi-turn reasoning and trustworthy tool integration.
Domain generalization approach for LiDAR semantic segmentation handling imperfect labels and sensor noise in autonomous driving scenarios.
RECODE agentic framework using code generation and derendering for verifiable visual reasoning on structured visuals with multimodal LLMs.
RL-100 real-world robotic manipulation framework combining diffusion visuomotor policies with imitation and reinforcement learning via clipped PPO.
Personalized collaborative learning framework with affinity-based variance reduction for heterogeneous multi-agent systems.
FALCON vision-language-action model incorporating 3D spatial foundation priors to improve reasoning and generalization in embodied AI tasks.
Interpretable operator-learning ML model for reconstructing electric field distributions from EFISH signal profiles in plasma physics.
Fairness-aware LoRA fine-tuning of vision-language models for medical imaging with differentiable MaxAccGap loss for demographic parity optimization.
ELERAG system enhancing retrieval-augmented generation with entity linking to improve factual accuracy in specialized domains like education.
ADHint method integrating difficulty-aware hints into reinforcement learning post-training to improve sample efficiency and reasoning generalization.
Theoretical analysis of distributed optimization with multiple local updates between communication rounds, proving acceleration guarantees.
Bayesian generative modeling framework for flexible conditional inference on arbitrary partitions of observed variables without fixed conditioning structure constraints.
Equivariant neural networks for robust object recognition under symmetric transformations and unusual viewing conditions.
FinTexTS dataset pairs financial time-series with semantic text data for multi-modal financial forecasting and analysis.
Analysis of performative chain-of-thought in reasoning models, showing models generate tokens without revealing internal beliefs via activation probing.
PolyBlocks: MLIR-based modular compiler infrastructure for AI programming frameworks and chips using affine analysis and analytical cost models.
VLN-Cache improves vision-language model inference efficiency for Vision-and-Language Navigation via semantic-aware token caching.
Megatron Core system optimizations for scaling Mixture-of-Experts model training across memory, communication, and computation constraints.
Covenant-72B: 72B parameter LLM trained via globally distributed, trustless peer-to-peer training over the internet without whitelisting.
Clinical feasibility study of AMIE, an LLM-based conversational AI for patient diagnostic history in real-world primary care workflows.
PostTrainBench benchmarks LLM agents' ability to automate post-training of language models, extending AI agents to AI research automation.