Retrieval-Conditioned Topology Selection with Provable Budget Conservation for Multi-Agent Code Generation
Multi-agent LLM code generation system selecting orchestration topology based on code structural complexity retrieval.
Multi-agent LLM code generation system selecting orchestration topology based on code structural complexity retrieval.
Mechanistic analysis of attention module roles in vision-language model decoders for architectural optimization.
Safety failures in chain-of-thought reasoning traces of large reasoning models with adaptive multi-principle steering mitigation.
Geometric analysis of LLM failure modes (conflict, hallucination) via parametric vs. working memory interactions.
Data attribution method for LLMs: FakeWiki benchmark to identify source documents supporting model responses.
Contrastive consistency model for generative graph prediction reducing inference cost vs. diffusion methods.
Saliency-aware regularization method for post-training quantization of large language models under memory constraints.
Budget control framework for LLM search agents optimizing tool calls and token allocation in multi-hop QA.
Knowledge-graph paths as intermediate supervision for self-evolving search agents, improving multi-step reasoning.
Analysis of reconstruction-concealment tradeoff in MLLM jailbreak attacks, studying safety mechanism vulnerabilities.
Investigation of linear decodability vs. correction gap in medical LLM failures via Overthinking behavior analysis.
Study of cross-component interference in LLM agent systems across 32 configurations, showing all-in approaches often degrade performance.
Multi-agent LLM framework (SAGE) for time-series anomaly detection using specialized analyzers for structured diagnosis.
SkillRet benchmark for skill retrieval in LLM agent systems, addressing critical challenge of selecting appropriate skills from large libraries under constraints.
SDFlow addresses exposure bias in autoregressive time-series generation via similarity-driven flow matching, reducing accumulated errors in long-horizon prediction.
ReFlect system enables LLM reasoning to detect and recover from errors across multi-stage long-horizon tasks, improving reliability beyond chain-of-thought approaches.
HyperLens analyzes LLM inference dynamics through fine-grained confidence trajectories, leveraging layer-wise magnification in transformers to quantify cognitive effort.
Transformer-based LLM for analog circuit design from natural language, addressing dataset scarcity and efficiency challenges in hardware design automation.
Graph-enhanced representation method for multi-sheet spreadsheet understanding to improve LLM-based data analysis agents' ability to handle heterogeneous schemas.
Proposes long-horizon Q-learning (LQL) to address error propagation in off-policy value-based reinforcement learning through n-step inequalities.
AGPO: Asymmetric group policy optimization for LLM reasoning verification in search ads relevance, addressing reasoning capability narrowing in RLVR.
Study on using language representations in auto-bidding for real-time advertising to explicitly control high-level intent and strategy.
Taklif.AI: LLM-powered platform automatically generating personalized college assignments based on student interests and abilities.
SANEmerg: Emergent communication framework enabling semantic-aware cooperation between heterogeneous AI agents in networking systems.
Risk intelligence system for Internet of Value using prediction engines and verification for composite risk in heterogeneous networks.
MLLM unlearning approach using null space constraints to remove target visual knowledge while preserving non-target knowledge in multimodal models.
User study on expert mathematicians interacting with AlphaEvolve evolutionary coding agent, characterizing intentmaking workflow for AI-guided discovery.
ICU-Bench: Benchmark for evaluating continual unlearning in multimodal large language models with sequential privacy deletion requests.
MAS-Algorithm: Multi-agent system workflow for solving algorithmic programming problems using structured reasoning and external tools.
Debiased knowledge tracing approach addressing selection bias in educational logging systems for student mastery estimation.
Prototype-based alignment method for heterogeneous federated learning across clients with different data distributions and model architectures.
TheraAgent: Agentic framework for iterative treatment planning using LLMs with verification loops instead of one-shot generation for safer, more comprehensive medical plans.
BehaviorGuard provides online backdoor defense for deep reinforcement learning using trigger-agnostic behavior analysis.
TACT mitigates agent drift in coding agents via activation steering, addressing overthinking and overacting failure modes in long-horizon tasks.
BioResearcher multi-agent system for translational medicine combining literature, trials, patents, and multi-omics analysis with auditable provenance.
Strat-LLM framework uses LLM agents for stock trading with real-time multi-source data integration and strategy alignment, tested live in 2025.
Novelty-based tree-of-thought search algorithm for improved LLM reasoning and planning with reduced token costs.
Visual fingerprinting method for analyzing how prompts, instructions, and parameters shape LLM output behaviors.
VibeServe: multi-agent agentic loop that automatically synthesizes bespoke LLM serving system stacks end-to-end.
Kernel embedding framework treating safety certification as classification problem for dynamical systems under uncertainty.
SPEED: layer-asymmetric key-value visibility policy for efficient long-context inference in decoder-only language models.
Framework for resource-constrained scheduling of agentic workflows under time and budget constraints with task dependencies.
CrossCult-KIBench benchmark for evaluating cross-cultural knowledge adaptation in multimodal language models.
Policy-guided model routing for cost-effective reasoning by dynamically routing chain-of-thought states across language models of different sizes.
LLM-based automatic heuristic design for combinatorial optimization using bottom-up paradigm from code to knowledge insights.
P-Guide: parameter-efficient method for classifier-free guidance in flow matching using single-pass inference with latent state modulation.
Skill1 framework for unified evolution of skill-augmented language model agents through reinforcement learning with persistent skill libraries.
Graphlets as structural vocabulary tokens for knowledge graph foundation models to enable discrete symbolic representation.
Policy invariance framework for testing reliability of LLM-as-a-Judge safety evaluation pipelines used in agent systems.
Post-Reasoning approach to improve LLM performance on tasks requiring minimal reasoning while reducing token consumption and latency.