PACED: Distillation and Self-Distillation at the Frontier of Student Competence
Framework for efficient LLM distillation that focuses training on problems at frontier of student capability.
Framework for efficient LLM distillation that focuses training on problems at frontier of student capability.
Protocol for detecting self-preservation behaviors in autonomous agents to distinguish intrinsic from instrumental objectives.
Framework coordinating multiple LLM-based agents through verification loop for complex query resolution with DAG decomposition.
Benchmark with 6,372 multimodal reasoning instances that evaluates LLM reasoning transparency through verifiable intermediate steps.
Research on semantic invariance property of LLM-based autonomous agents under input variations to ensure stable reasoning.
Method using counterfactual thinking to identify and address bias and fairness issues in machine learning models.
Research on automated prompt generation and optimization techniques for improving LLM performance through meta-prompting approaches.
Ayn: domain-specific tiny language model pretrained from scratch for Indian legal NLP tasks as alternative to large LLMs.
Survey examining computerized adaptive testing through machine learning lens, covering personalized assessment methods across domains.
Universal approximation theorem and operator learning methods for continuous nonlinear operators in Banach spaces using orthogonal projections.
TraffiDent dataset aligning traffic dynamics and incident data across 16,972 nodes for understanding their interplay.
Analysis of how skip connections in deep networks enhance adversarial example transferability across models.
Time series forecasting approach accounting for latent confounders using causal inference to improve prediction accuracy.
Causal inference method using LLMs to quantify effects of textual interventions on social systems from observational data.
VisionZip reduces computational costs in vision-language models by compressing redundant visual tokens while maintaining performance.
RRNCO addresses real-world deployment of neural combinatorial optimization for vehicle routing by handling asymmetric costs and edge-based features.
Review of LLM-driven approaches for creating virtual agents with personality in VR environments using multimodal outputs.
Mask Fine-Tuning (MFT) introduces a novel LLM fine-tuning method that improves performance by selectively masking model components without updating weights.
MegaScale-Data addresses computational challenges in training large foundation models from multiple data sources by optimizing dataloader distribution across parallel ranks.
Credit assignment method (QLLM) for multi-agent RL eliminating predefined mixing networks through improved value decomposition and interpretability.
Nemotron-CrossThink extends RL-based self-learning from math reasoning to broader domains using verifiable reward structures and diverse tasks.
PCCL library for performant collective communication in distributed AI training on GPU supercomputers, addressing NCCL limitations.
Aitomia platform combining LLM-based agents and chatbots to assist with atomistic and quantum chemical simulations setup and analysis.
VideoSafetyEval benchmark with 11.4k video-query pairs across 19 risk categories for evaluating and defending Video LLM safety.
Method for improving LLM reasoning without expensive RL or high-quality demonstrations using weak supervision and incentive signals.
Inference-time alignment method for LLMs that searches in continuous response space using reward models for improved exploration.
SVD-based compression method (ERC-SVD) for efficient LLM deployment with error control and low-rank approximation techniques.
Analysis of implicit regularization in overparametrized deep neural networks and improved out-of-distribution generalization via variational methods.
Dynamic benchmark framework (NetArena) for evaluating AI agents in network automation with production-level complexity and reduced contamination risk.
Adaptive multi-objective reinforcement learning method for balancing exploration and skill diversity in skill-based RL pretraining.
Benchmark for evaluating multimodal LLM-based front-end code generation with modern development frameworks and evaluation metrics.
Curriculum learning approach scheduling tasks from easy to hard to improve LLM reasoning via reinforcement learning, inspired by DeepSeek-R1.
BIS Reasoning 1.0: Japanese benchmark with 1K+ syllogistic problems evaluating belief bias and inconsistent reasoning in LLMs.
AVA-Bench: systematic evaluation benchmark for vision foundation models addressing blind spots in VQA evaluation protocols.
TRACED: unsupervised environment design using regret approximation for co-learning to improve deep RL agent generalization.
Rationale-Enhanced Decoding improves chain-of-thought prompting in vision-language models by optimizing intermediate reasoning generation.
Lumos-1: LLM-based autoregressive video generation using discrete diffusion with efficient architecture avoiding external encoders.
SOAR: self-improving method integrating language models into evolutionary program synthesis for challenging tasks like ARC-AGI.
FingerTip 20K: benchmark for proactive mobile LLM agents with 20K tasks, evaluating multimodal agents using contextual data without explicit instructions.
Neural Combinatorial Optimization solver for min-max heterogeneous vehicle routing with multiple vehicles using novel decoding approach.
EvolvR: self-evolving method for story evaluation using LLM-as-judge with pairwise reasoning to improve generation guidance.
Novel benchmarking system evaluating LLM-based agent capabilities for single-cell omics data analysis, assessing planning and code generation.
Systematic study of post-training quantization methods for diffusion LLMs to enable edge device deployment, comparing compression techniques.
UTRL: reinforcement learning framework training LLMs to generate high-quality unit tests automatically, addressing test generation challenges.
Research evaluating Law-Following AI framework for embedding legal compliance in advanced AI agents, analyzing legal personhood constructs and technical feasibility.
Reinforcement learning approach for radiology report generation using FactScore-based rewards with reduced data requirements.
Framework evaluating robustness of Vision-Language-Action models under real-world physical variations for robotic tasks.
Matched-compute study evaluating synthetic data interventions for in-context learning in language models. Tests mechanism-targeted pretraining effects.
Method for reducing LLM agent inference costs through trajectory reduction. Addresses token cost efficiency in multi-turn agent systems for software engineering.
Technique reducing LLM reasoning model overthinking through decoupled rewards and curriculum scheduling. Addresses excessive token generation without performance gain.