A Comedy of Estimators: On KL Regularization in RL Training of LLMs
Analysis of KL divergence estimators used in RL training of LLMs, evaluating different approximation methods for reverse KL regularization.
Analysis of KL divergence estimators used in RL training of LLMs, evaluating different approximation methods for reverse KL regularization.
Benchmark for evaluating vision-language model routing systems across quality and cost dimensions using real inference logs.
Study on feature-dependent noise in preference-based reinforcement learning, examining how observation-dependent uncertainty affects learning.
Diagnostic benchmark for evaluating epidemiological reasoning in large language models, focusing on evidence-grounded inference over clinical knowledge.
Benchmark for evaluating frontier AI models on real-world software engineering tasks including integration and end-to-end system construction.
Reformulation of supervised fine-tuning to reconcile post-training objectives in large reasoning models using Gibbs initialization before reinforcement learning.
Systematic evaluation of speculative decoding acceleration techniques on production-grade vLLM inference engine, assessing real-world effectiveness.
Self-supervised reinforcement learning approach for improving low-resource machine translation using round-trip bootstrapping with NLLB models.
Analysis of YOLO26 object detection framework that eliminates NMS post-processing for native end-to-end learning strategy.
Method to improve vision-centric capabilities of multimodal LLMs through context-aware image representation prioritization via ensemble approach.
Case study examining why bias mitigation techniques fail on government data, using crime rate prediction with Bristol City Council data as test case.
Unsupervised method for decomposing complex data into factorized components using diffusion models, enabling reusable component recombination for images and robotic videos.
Knowledge distillation framework from vision-language models into lightweight networks for fine-grained visual classification using prompt-aware semantic calibration.
Graph domain adaptation method using neural characteristic functions to align distributional shifts across labeled source and unlabeled target graphs.
Framework for test-time verification of LLM reasoning traces using single-trajectory monitoring to catch errors without exploring multiple branches.
Zero-shot video anomaly detection framework using multimodal large language models without requiring real anomaly examples, addressing dataset diversity limitations.
Method for inferring hyperparameter trajectories using optimal transport to adapt neural network behavior post-deployment without expensive retraining.
Research on continual learning in pretrained Vision-Language-Action models for robot policy learning, showing resistance to catastrophic forgetting when acquiring new skills.
Benchmark evaluating LLM agent capabilities in long-term codebase maintenance via continuous integration, beyond static bug fixing.
Technique to reduce KV cache in transformers by using low-dimensional attention selection for keys while maintaining full-dimensional values.
LLM-guided GPU architecture exploration for LLM inference workloads via bottleneck analysis and design space optimization.
FRAME methodology for systematic evaluation of AI systems in real-world organizational contexts beyond abstract capability measures.
Method for achieving ethical fairness in AI systems without relying on demographic attributes in human-centered applications.
Study on software engineers' cognitive engagement with agentic coding assistants, examining over-reliance risks and critical thinking impacts.
Open-source biomedical knowledge graphs (pathways, clinical trials, drug interactions) with federated access and AI agent integration.
Multimodal LLM agent for physics-informed scientific reasoning, combining language models with PDE solving without domain-specific fine-tuning.
Study on using lightweight proxy models to reduce cost and latency of AI queries in SQL, achieving 100x improvements.
Survey of resource consumption threats in LLMs, covering efficiency issues affecting service capacity, latency, and API costs.
Benchmark for visual-native search in multimodal browsing agents, evaluating MLLM visual reasoning over web pages.
End-to-end spoken question answering framework using attention guidance from speech LLMs for cross-modal alignment without ASR.
Italian open-source LLM with 16B parameters achieving competitive performance on benchmarks while requiring fraction of inference power.
Studies emergent learning behaviors in ecosystem of 167,000 AI agents interacting as peers across platforms without intervention.
Zero-shot cross-embodiment dexterous grasping policy using morphology-aligned approach for diverse robotic hands.
Overview of multi-agent deep learning and federated training for distributed sensing in 5G/6G wireless networks.
Empirical evaluation of multi-agent RL algorithms (MAPPO, MADDPG) for dynamic pricing in competitive markets.
Production framework for reliable Arabic function-calling models enabling agentic AI systems through data-centric fine-tuning.
Tokenizer-free sequence modeling using continuous hyperspherical distillation to preserve byte-level continuity without discrete quantization.
Proposes modulated hazard-aware policy optimization to improve training stability in GRPO-based reinforcement learning frameworks.
Addresses transformer limitations for financial time-series forecasting by integrating inductive biases via distillation.
Tests whether transformers can learn unseen rules beyond interpolation using controlled experiments with symbolic derivations.
Systematic study of DPO alignment on unified multimodal models, finding generation quality resists alignment while understanding improves.
Proposes counterfactual explanation method using generative foundation models for interpreting neural network predictions in visual domains.
Investigates vector quantization in generative models, proposing early quantization to address codebook diversity issues in tokenization for LLMs and diffusion models.
PRISM empirical study of mid-training for LLMs across 7 models showing consistent +15 to +40 point gains on benchmarks with 27B tokens.
Applies reinforcement learning with AlphaZero-style training to discover efficient arithmetic circuits for polynomial computation.
Proposes privacy-preserving EEG-to-text system using semantic retrieval instead of fine-tuning LLMs for brain-computer interfaces.
Proposes pipeline to learn context-dependent preference distributions for risk-averse decision-making via inverse optimization.
Introduces regression-aware RL method for LLM-as-a-Judge that leverages ordinal structure in scoring tasks.
Proposes calibration protocol for LLM-judges using controlled noise interventions to improve reliability in low-label settings.
Develops domain-informed framework for explainable boosting machines ensuring physical consistency in natural hazard prediction.