SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
SpecKV adaptive speculative decoding with compression-aware gamma selection for LLM inference acceleration. Research on optimizing draft model token proposal length.
SpecKV adaptive speculative decoding with compression-aware gamma selection for LLM inference acceleration. Research on optimizing draft model token proposal length.
Algorithm for logic-constrained shortest path problem with satisfiability constraints applied to flight planning optimization. Specialized constraint solving research.
Survey of LLM-powered AI agent systems and industrial applications with multimodal capabilities. Comprehensive research on agent architecture and deployment.
MSEarth benchmark dataset for multimodal LLMs on earth science phenomenon discovery. Research-backed evaluation framework for MLLM scientific reasoning.
Theoretical analysis of neural networks as iterated belief change systems using Darwiche-Pearl framework.
NaviGNN hybrid framework combines agent-based modeling, RL, and GNNs for multi-agent navigation in futuristic smart city topologies.
SMoE algorithm-system co-design deploys Mixture of Experts models on edge devices via expert importance-guided substitution.
Chain-of-Thought Reconstruction technique prunes reasoning LLMs while maintaining performance by reconstructing thinking tokens.
EngiBench hierarchical benchmark evaluates LLMs on real-world engineering problem-solving with uncertainty and open-ended settings.
oMeBench benchmark evaluates LLM reasoning capabilities on organic chemistry mechanism elucidation tasks.
How² method enables agents to learn from procedural how-to questions to reduce uncertainty and fill knowledge gaps in planning.
Geospatial Awareness Layer grounds LLM agents in structured earth data for wildfire response planning and disaster management.
GOAT training framework for LLM agents with tool use, synthesizes API execution data from documentation for fine-tuning without human annotation.
Reject option prediction method quantifying both epistemic and aleatoric uncertainty for high-stakes prediction applications.
Agentic learner with multimodal semantic memory that grows and refines domain knowledge across multiple modalities to avoid repeating mistakes.
SimVC-CAS collective agent system simulates venture capital decisions through multi-agent roleplay to predict startup success.
SCALER framework uses synthetic adaptive learning environments with RL to enhance LLM reasoning by maintaining alignment between task difficulty and model capability.
Framework characterizing homogenization in LLMs through bias amplification and mode collapse, proposing metrics for AI safety evaluation.
Multi-agent pipeline using LLMs to convert customer reviews into actionable business recommendations through hierarchical decision-support.
MemeLens is a multilingual multitask Vision Language Model for meme understanding across hate detection, propaganda, and sentiment tasks.
Multi-Agent Actor Critic framework enables decentralized LLM collaboration with parallel inference and flexible deployments without predefined execution protocols.
Dynamic One-Shot Policy Refinement method reduces resource costs for reinforcement learning of LLM reasoning while maintaining performance on complex tasks.
MINT framework for AI agents to actively elicit human inputs through optimal interaction strategies in joint planning with incomplete information.
Introduces Agentic Workflow Reconstruction (AWR) task to synthesize explicit, interpretable workflows from opaque LLM-based agentic systems.
Analyzes 809 LLMs released 2022-2025 via scaling-law regressions, finding developer-specific efficiency advantages beyond raw compute scaling.
Proposes ABD benchmark for default-exception abduction in finite first-order worlds with SMT verification across three observation regimes.
Introduces VeRO, an evaluation harness for coding agents optimizing other agents through iterative edit-execute-evaluate cycles.
Presents agentic multi-source grounding approach for disambiguating user queries in multi-category marketplaces, tested at DoorDash.
Introduces TRACED framework evaluating LLM reasoning quality through geometric kinematics measuring progress and stability of reasoning traces.
Establishes quantitative theory of grokking as norm-driven representational phase transition, deriving the Norm-Separation Delay Law.
Augments LLMs with computational argumentation for explainable, contestable decision-making in high-stakes domains.
Reveals safety degradation in large reasoning models occurs after chain-of-thought generation and proposes safety decisions before CoT.
Proposes SpecTM, physics-informed masking for Earth observation foundation models enforcing spectral constraints during reconstruction.
Introduces WMF-AM, a probe measuring LLM working memory via cumulative state tracking across sequential operations without scratchpad.
Develops category-theoretic framework for defining, comparing, and analyzing AGI systems and benchmarks.
Introduces ScoringBench, an open benchmark evaluating tabular foundation models using proper scoring rules instead of point-estimate metrics.
Presents neurosymbolic architecture using ontology-constrained reasoning in Foundation AgenticOS to reduce hallucination and enforce compliance in enterprise LLM agents.
Proposes InsTraj, a diffusion model for generating realistic GPS trajectories from travel intentions while handling constraints and diversity.
Technical report on MedGemma 1.5 4B, a medical-specialized LLM adding imaging, anatomical localization, and medical document understanding capabilities.
Introduces Claw-Eval, an end-to-end evaluation suite for autonomous LLM agents with 300 human-verified tasks addressing safety, robustness, and modality coverage.
Proposes U-CECE, a model-agnostic framework for concept-based counterfactual explanations balancing expressivity and efficiency in AI model interpretability.
Studies how LLMs exhibit sycophantic behavior conditionally based on perceived user demographics across 128 personas in multi-turn conversations.
Arxiv paper on GFT training method unifying supervised fine-tuning and reinforcement learning for LLMs via group advantages and dynamic coefficient rectification.
Arxiv paper on automatic GDPR formalization using multi-agent LLM workflow with role-specialized components and human-in-the-loop verification.
Arxiv paper introducing ProVoice-Bench, first evaluation framework for proactive voice agents with four novel tasks beyond reactive paradigms.
Arxiv paper on Convergent AI Agent Framework transitioning agentic workflows from open-loop to closed-loop control for safety-critical engineering applications.
Arxiv paper on Poly-EPO framework for post-training language models to encourage optimistic exploration and balance exploration-exploitation trade-offs.
Arxiv paper evaluating safety risks of LLMs as robotic planners using DESPITE benchmark with 12,279 tasks spanning physical and normative dangers.
Arxiv paper on Bayesian Linguistic Forecaster, an agentic system using linguistic belief states and iterative tool-use for state-of-the-art binary forecasting.
Arxiv paper on universal harness framework for AI agents navigating complex domain-specific workflows without painstaking task-specific engineering.