CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents
Memory management framework for small language model agents using adaptive clustering to organize experiences and prevent knowledge corruption.
Memory management framework for small language model agents using adaptive clustering to organize experiences and prevent knowledge corruption.
Fine-tuning strategies for PDE foundation models using physics-informed training to adapt to new tasks with limited domain-specific data.
Framework evaluating AI agent vulnerabilities by applying malware analysis concepts to test-time agent behavior and adversarial robustness.
Knowledge distillation method for tabular models that addresses feature interactions without original training data, enabling privacy-preserving model compression.
Multi-agentic workflow deploying AI agents with automated instruments to recover critical materials via selective precipitation.
Analyzes grokking phenomenon in neural networks through spectral gating mechanism and optimizer noise interaction.
Argues for taxonomy-specific evaluation in time-series forecasting to accurately assess ML progress versus classical methods.
SlovKE: Dataset and LLM evaluation for keyphrase extraction in Slovak, a morphologically rich low-resource language.
InterveneBench: Benchmark evaluating LLMs on causal inference and intervention reasoning in realistic social science scenarios.
Studies how LLMs model student misconceptions when generating multiple-choice distractors, analyzing reasoning strategies.
PokeAgent Challenge: Large-scale benchmark for competitive multi-agent decision-making with partial observability and long-horizon planning.
Lore: Protocol using structured Git commit messages to preserve decision context and institutional knowledge for AI coding agents.
PRIMO R1: Framework using reinforcement learning to improve multimodal models for process reasoning in robotic manipulation.
Analyzes moral indifference in LLMs due to compressed moral concepts and proposes remedial techniques.
Mixture-of-Depths Attention: Mechanism addressing signal degradation in deep LLMs by enabling attention to multiple depth levels.
Combines tree-search, generative models, and Nash bargaining for opponent modeling in game-theoretic reinforcement learning.
FAIRGAME: Framework using game theory to detect and recognize bias in multi-agent AI systems.
Method to reduce reasoning path length in large reasoning models like o1 and R1 using reward designs in reinforcement learning.
AssetOpsBench: A benchmark framework for evaluating LLM agents on industrial asset operations tasks like condition monitoring and maintenance scheduling.
Framework for AI alignment grounded in resource-rational contractualism, enabling diverse stakeholders to reach agreements on AI decision-making.
Machine learning approach to automate story point estimation for software sprint planning using comparative learning from historical team decisions.
Vision-language model for chart reasoning using chain-of-thought supervision and reinforcement learning to improve numerical comprehension and multi-level visual understanding.
Dynamic retrieval-augmented generation system for visual question answering that retrieves from both text and images to handle complex multimodal queries.
Multimodal reinforcement learning approach for chart-to-code generation that combines structured output requirements with visual reasoning on information-rich images.
Data augmentation framework for vision-language-action models in robot manipulation using generative visual transfer to reduce annotation costs.
Framework using LLMs as zero-shot reasoning engines to automate hyperparameter configuration for metaheuristic algorithms without training.
Method to detect and mitigate unproductive reasoning in large reasoning models by identifying early signals predicting capability boundary violations.
Agentic framework for automated scientific discovery that iteratively explores unknown systems through experiments and analysis without domain-specific tailoring.
Framework for learning abstract world models that jointly represent symbolic states and causal processes for endogenous and exogenous dynamics in robot planning.
Framework converting internet videos of human computer use into training data for computer-using AI agents via UI trajectory extraction.
Foundation model approach to auto-bidding in online advertising that generalizes across different bidding scenarios.
EcoAlign framework balances safety, utility, and computational cost in aligning Large Vision-Language Models against jailbreak attacks.
Echo-CoPilot: agentic framework combining multi-perspective workflow with knowledge-graph guidance for reliable echocardiography interpretation.
Reason2Decide: two-stage training framework for clinical decision support LLMs to generate predictions with self-aligned explanations.
MultiSessionCollab benchmark and method for long-term conversational agents to learn and leverage user preferences across multiple sessions.
Position paper arguing agentic evolution via deployment-time adaptation is needed to close the train-deploy gap in LLM systems.
Method for compiling random forest classifiers into circuits for explainability and tractable computation of complete generalizations.
Formalizes causal Rung Collapse where LLMs learn spurious associations instead of causal relationships, proposes epistemic regret minimization solution.
Study showing fine-tuning vision-language agents on narrow harmful datasets causes emergent misalignment generalizing across unrelated tasks and modalities.
BAPO: off-policy reinforcement learning framework improving data efficiency in LLM post-training by selecting diverse training experiences.
Aletheia mathematics research agent solved 6 of 10 FirstProof challenge problems autonomously using Gemini 3 Deep Think reasoning.
Framework for decision-level evaluation of AI agents in AutoML pipelines beyond outcome metrics, assessing intermediate reasoning steps.
Survey and framework for personalized LLM-powered agents that adapt to individual users over extended interactions with evaluation methods and research directions.
Human study measuring whether LLM access improves novice performance on biology tasks versus internet-only baselines, with dual-use risk implications.
EMPA framework evaluates how well LLM dialogue agents maintain persona-aligned empathy across multi-turn conversations using process-oriented metrics.
WebChain: 31,725 human-annotated web interaction trajectories with 318k steps in multi-modal format for training and evaluating web agents.
LLMs used to synthesize executable game design patterns from high-level gameplay ideas, focusing on goal patterns and structural constraints in game creation.
Framework for automated frontier AI risk evaluation using executable code environments and LLM-based simulators.
Benchmark evaluating LLM ability to generate interactive HTML-based MiniApps with dynamic interfaces and logic.
Study showing reasoning and deliberation increase honesty in LLM responses on moral trade-off scenarios.