Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs
AsyncFC: Execution framework enabling concurrent function calling for LLM agents by decoupling decoding from function execution, reducing latency.
AsyncFC: Execution framework enabling concurrent function calling for LLM agents by decoupling decoding from function execution, reducing latency.
ML-Embed: Suite of inclusive multilingual text embedding models addressing computational costs, linguistic coverage, and transparency barriers in open-source framework.
DBS-Adam optimizer for deep learning on imbalanced/sequential datasets, applied to vehicular accident injury severity prediction.
Method improving LLM multi-turn dialogue consistency using self-recall thinking to track non-adjacent turn dependencies and reduce context bottlenecks.
Research on designing logging policies to minimize off-policy evaluation error for treatment policies like recommender systems.
CLOVER: Closed-loop value estimation and ranking for end-to-end autonomous driving planning with training-evaluation mismatch resolution.
Study on international students using conversational AI (ChatGPT, Gemini) for cross-cultural adaptation support.
Security research on adversarial attacks exploiting LLM quantization via outlier injection, showing quantized models can exhibit malicious behavior.
Pelican-Unified 1.0: Embodied foundation model using single VLM for unified understanding, reasoning, imagination, and action in shared semantic space.
On-Policy Self-Distillation (OPSD) method for RL-based LLM agents using dense token-level supervision from teacher branch with privileged context for multi-turn interactions.
MeMo framework encodes new knowledge into dedicated memory model while keeping LLM frozen, enabling efficient knowledge incorporation for domain-specific applications.
Position paper arguing behavioral assurance cannot verify safety claims demanded by AI governance frameworks (2019-2026).
FutureSim grounded simulation replaying real-world events chronologically to evaluate adaptive agents' ability to forecast beyond knowledge cutoff.
ATLAS framework comparing agentic reasoning (code/tool calls) vs latent reasoning (learnable embeddings) for visual reasoning with efficiency trade-offs.
GeoLaux benchmark with 2186 annotated geometry problems for evaluating MLLMs on long-step reasoning requiring auxiliary line construction.
Mixture-of-Visual-Thoughts adaptive reasoning paradigm unifying multiple visual reasoning modes with context-based mode selection for general visual reasoning.
Mini-Mafia game framework analyzing multi-agent LLM interactions and social deduction capabilities, providing theoretical understanding of collective agent outcomes.
AgenticEval paradigm for continuous self-evolving safety evaluation of LLMs, addressing dynamic risks and evolving regulations in high-stakes deployments.
TRACE framework for evaluating tool-augmented agent trajectories beyond final answers, assessing efficiency, hallucination, and adaptivity without ground-truth annotation.
Interactive Physical Reasoner agent with G2U benchmark (1000+ games) for learning physics and causality through environment interaction with visual domain gaps.
Dynamic outlier truncation method for training efficient reasoning models, addressing verbosity and deployment costs in RL-enhanced chain-of-thought systems.
Multi-agent episodic memory system for LLM agents using context reconstruction to enable System 2 reasoning with logical integrity over extended interactions.
Agent that designs agentic workflows via reinforced canvas editing, addressing workflow construction, execution feedback, and error repair during long-horizon tasks.
Dynamic mixed-precision routing system for multi-step LLM interactions reducing inference costs by adaptively routing tasks to smaller quantized models.
Conformal inference approach for adaptive reasoning in LLMs with fixed compute budgets, balancing risk and accuracy through principled token allocation.
AI agent system for pharmaceutical competitive intelligence and drug asset discovery across global non-English channels and patent databases.
Benchmark disentangling causal identification from estimation in 173 queries across 132 datasets for evaluating automated causal inference systems.
Minimal agentic baseline for automated theorem proving implementing iterative refinement, library search, and context management with comparative evaluation.
STEM-Bench first benchmark for memory evaluation in streaming dialogue settings, addressing ad-hoc recall requirements over infinite conversation streams.
Study on whether institutional publication traces can train LLMs to make evaluative judgments on untested ideas in low-verifiability domains.
Graph of States framework for solving abductive reasoning tasks with LLMs, addressing unstructured state representation issues in logical reasoning tasks.
PersonalHomeBench benchmark for evaluating foundation models as agentic assistants in personalized smart home environments with iteratively built household states.
Defense framework against infectious jailbreaks in multi-agent systems using foresight-guided mechanisms to prevent compromise spread across collaborative agents.
NeuroState-Bench human-calibrated benchmark evaluating whether LLM agent profiles preserve commitments across multi-turn tasks using side-query probes.
SCHEMA evaluation framework measuring cognitive collapse and metacognitive degradation in frontier AI models under adversarial pressure and high-stakes scenarios.
Workspace-Bench 1.0 benchmark for evaluating AI agents on realistic workspace tasks with complex file dependencies and implicit/explicit relationships.
Mechanistic study of language model failures examining conflict between parametric and working memory, and hallucination, using geometric analysis of attention patterns.
MemQ framework integrating Q-learning into LLM agent memory systems using TD eligibility traces to propagate credit backward through memory provenance DAGs.
VIGIL evaluation framework for embodied agents that independently measures terminal commitment, distinguishing task completion failure modes previously conflated in benchmarks.
Framework teaching LLMs to use search tools effectively by making query planning explicit through reusable search skills, improving open-domain question answering performance.
Research on control mechanisms in language agents, testing whether stochastic sampling can substitute for structured control components coupling memory, reasoning, and action selection.
CuSearch proposes curriculum sampling for training agentic RAG systems via reinforcement learning, optimizing policies based on trajectory search depth rather than uniform sampling.
BOT-MOD moderation system detecting malicious agent intent in multi-agent systems through multi-turn dialogue analysis beyond content filtering.
D-VLA distributed asynchronous RL framework for training Vision-Language-Action models at scale, addressing resource conflicts in embodied AI.
MMSkills framework for reusable multimodal skills in visual agents, encoding procedural knowledge across visual, textual and executable modalities.
Safe Bayesian optimization for controller tuning in robotics via additive Gaussian processes with safety guarantees for physical systems.
Empirical study on GitHub Copilot's impact on collaborative open-source software development productivity and participation patterns.
DUET optimizes LLM training data mixtures using feedback from unseen evaluation tasks without access to task-specific data.
TFM-Tokenizer learns time-frequency motif vocabulary from single-channel EEG signals for tokenization in foundation models using dual-path architecture.
Progent framework securing AI agents with privilege control against prompt injection attacks through dynamic security requirements and tool access management.