MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
MAD-OPD uses multi-agent debate to overcome single-teacher capability ceiling in on-policy distillation for agentic tasks.
MAD-OPD uses multi-agent debate to overcome single-teacher capability ceiling in on-policy distillation for agentic tasks.
Self-contrast decoding strategy for diffusion large language models to improve generation quality via high-information-density context modeling.
Empirical study of how developers use LLMs in software design tasks, surveying practitioners on benefits and drawbacks.
LiveFMBench: systematic study of LLM and agent-based formal specification generation for C programs with contamination awareness.
Verbal-R3 bridges retrieval and LLM reasoning via verbal annotations—analytic narratives connecting queries to retrieved contexts in RAG systems.
Framework for capturing expert cognition in practice-based domains to enhance AI-driven educational systems.
Medmarks: open-source benchmark suite with 30 benchmarks for evaluating LLMs on medical tasks including QA, information extraction, and calculations.
HepScript DSL enabling human-AI collaborative data analysis for high-energy physics workflows using agentic LLMs and domain-specific knowledge.
Theoretical analysis of generalization properties in multimodal metric learning with incomplete or redundant data.
Analysis of adversarial attacks on vision-language models, distinguishing between output perturbation and precise injection attacks.
Multi-agent autonomous testing system using LLM with LangGraph orchestration for UI test repair in enterprise applications.
FT-RAG: fine-grained retrieval-augmented generation framework for complex table reasoning using entry-level decomposition and semantic comprehension.
ProMORNA: multi-objective reinforcement learning framework for designing full-length therapeutic mRNA from protein sequences balancing stability and safety.
Advances machine learning surrogates for CFD using GNNs and transformers, addressing training paradigm bottlenecks in physics simulations.
Proposes agentic AI-native 6G networks where LLM-based agents operate as bounded, policy-governed entities for autonomous intelligence.
Autonomous multiagent framework automating mechanistic interpretability of LLMs through iterative hypothesis testing and feature discovery agents.
Neuro-symbolic agents framework combining LLMs with formal methods to enable flexible requirements reuse while maintaining structural validity and consistency.
Thesis on model merging paradigm: combining independently trained neural networks in weight space without optimization or original training data access.
SkillGraph-Service: hybrid microservice merging knowledge graphs with LLM fallback for competency framework interoperability and skill search.
Introduces S²R², a segment-level robustness framework for LoRA-tuned language models that enforces consistency at entity and relation levels beyond whole-sequence.
Tests causal inner products for cross-lingual concept transport in transformer representations across 17 models, finding anti-concentration phenomena.
Examines how agentic AI interfaces should prioritize explanation and oversight communication over routine user interaction during workflow execution.
Introduces Prosa, a Brazilian Portuguese LLM evaluation benchmark using rubric-based scoring with multi-judge filtering to reduce judge model bias.
Studies AI alignment through law-and-economics models of deterrence, analyzing how agentic AI systems respond strategically to incentive structures.
Brain-inspired spiking neural network with oscillatory dynamics and time-delayed coordination for cognitive learning mechanisms.
TRIMMER uses self-supervised reinforcement learning for domain-agnostic video summarization without manual annotations.
IMPACT-HOI mixed-initiative framework for egocentric video annotation of Human-Object Interactions to generate robot manipulation training data.
IMPACT-Scribe framework for dense video annotation using interactive correction-driven collaboration between human annotators and models.
MissBGM Bayesian generative model for missing data imputation with uncertainty quantification using neural networks and Bayesian inference.
Adaptive differential privacy method for fall detection in sensor data that applies class-aware noise to preserve prediction performance.
GRAVITY architecture for long-horizon conversational agents that structures retrieved memory fragments with relational and temporal anchors.
LLM-based adaptive agent for extracting information from BIM models through iterative exploration instead of static query translation.
Technique to surgically remove memorization signatures in unlearned LLMs using cross-sequence probes without capability loss.
SplitZip compression method for KV cache transfer in disaggregated LLM serving, optimizing prefill-decode architecture bottleneck.
TCDA model for conversational sentiment analysis using graph networks and positional encoding to handle multi-turn dialogues with temporal context.
Training-free method to mitigate object hallucination in vision-language models by steering caption generation.
Security analysis of agentic-AI runtimes showing critical gaps in action auditing and safety properties across tool calls and device actuation.
Research on dynamic grounding failures and repair mechanisms in multi-agent LLM negotiation across conversational turns.
Study of the Compliance Gap where AI agents verbally agree to instructions but fail to follow them, examining the disconnect between stated and actual behavior.
LLM-guided evolutionary search (AlphaEvolve) to discover closed-form solutions for radar power allocation in multi-target tracking.
Selector-guided curriculum learning approach for one-shot reinforcement learning from verifiable rewards in LLM math reasoning.
Benchmark (RMGAP) for evaluating reward model generalization across diverse user preferences in RLHF-aligned language models.
Framework for remote control with communication-constrained channels where actors receive action guidance from a centralized controller.
Evaluation of data-poisoning backdoor attacks on contrastive learning models and their generalization across datasets.
Analysis of hidden-state dynamics in large reasoning models to determine if extended reasoning traces reflect genuine internal computation or verbosity.
Research on exploration budget allocation in cooperative multi-agent reinforcement learning using intrinsic motivation and novelty bonuses.
Method for selecting optimal training subsets under label noise by leveraging data symmetries, detecting low-quality examples via k-NN based analysis.
Chart-FR1: Multimodal LLM benchmark for fine-grained reasoning on high-information-density charts with multiple subplots and dense annotations.
SANTA: Sparse attention method for memory-bound LLM inference that samples from post-softmax distribution to reduce KV cache bandwidth at long contexts.
RefusalGuard: Geometry-preserving fine-tuning method maintaining safety alignment in LLMs by analyzing activation-space feature changes during task-specific adaptation.