Regret-Guided Search Control for Efficient Learning in AlphaZero
Search control method for AlphaZero using regret-guided learning, improving sample efficiency by restarting from valuable game states rather than initial positions.
Search control method for AlphaZero using regret-guided learning, improving sample efficiency by restarting from valuable game states rather than initial positions.
Systematization of knowledge on agentic skills—reusable procedural capabilities with explicit conditions, policies, and interfaces for reliable long-horizon workflows in LLM agents.
Airavat: agentic framework automating internet measurement workflows and verification against methodological standards using AI agents for tool orchestration.
Research on efficient reasoning in LLMs using Chain-of-Thought, addressing computational overhead through reward shaping and reinforcement learning optimization.
Agentic data synthesis approach to help vision-language models and diffusion models detect and mitigate visual artifacts in AI-generated images.
Theoretical analysis connecting law of robustness to robust generalization, resolving open problem on overparameterization requirements for robust interpolation.
Vision paper proposing Agentic Infused Software Ecosystem (AISE) framework rethinking software development tools and practices for autonomous AI agents.
CrystaL method enables spontaneous emergence of visual latent representations in MLLMs through improved supervision for latent chain-of-thought reasoning.
Report-supervised learning for brain lesion segmentation using multimodal MRI and radiology report constraints on substructures.
MIP Candy is a modular PyTorch framework for medical image processing with support for volumetric data, multiple formats, and domain-specific training.
VAUQ provides vision-aware uncertainty quantification for evaluating Large Vision-Language Model outputs and reducing hallucinations in vision-conditioned tasks.
Localized domain adaptation method for offline reinforcement learning addressing dynamics mismatch between source and target domains.
Analyzes how graph topology interacts with GNN learned preferences through massive activations, investigating oversmoothing and oversquashing artifacts.
Multi-agent deep RL system for training cooperative-competitive robot teams in simulation with transfer to real-world robotic applications.
Study examines security vulnerabilities in LLM-driven agents through Agent-Mediated Deception attacks, revealing human perception susceptibility to compromised copilots.
SparkMe uses LLMs to automate semi-structured qualitative interviews with adaptive topic exploration balanced against systematic coverage.
PVminer uses ML/NLP to detect patient voice in healthcare text data, scaling qualitative analysis of patient-generated messages across health systems.
XMorph framework combines deep learning with LLM assistance for interpretable brain tumor classification and segmentation on medical images with computational efficiency.
Study reveals that optimizing Pass@k metric for LLM code generation can degrade Pass@1 performance due to prompt interference, analyzing trade-offs in inference-aware fine-tuning methods.
Reflective Test-Time Planning enables embodied LLMs to learn from failures through in-action and post-action reflection modes, improving robot task reasoning and reducing repeated mistakes.
Analysis revealing test-time training with KV binding functions as linear attention rather than memorization-based meta-learning.
Comprehensive survey on optimization techniques for LLM-based agents covering prompt design, fine-tuning, and performance improvements.
STPR framework uses LLMs to generate formal constraints for robot navigation from complex natural language spatial and conditional specifications.
Method enabling LLMs to control embodied agents by generating and iteratively refining control policies from natural language descriptions.
Programming by Backprop enables LLMs to learn procedural knowledge from declarative instructions during training, improving few-shot capabilities.
AutoEDA applies microservice-based LLM agents to automate Electronic Design Automation workflows, replacing Tcl scripting with natural language.
Comprehensive analysis of massive activation patterns in transformer training using Pythia model family with public dataset release.
TASER uses AI agents to extract structured data from unformatted financial tables for QA and data normalization tasks.
Hybrid Deep Searcher combining parallel and sequential search reasoning for large reasoning models with RAG enabling test-time scaling.
DS-STAR: Data science agent for multi-source data integration and exploration synthesizing open-ended insights across heterogeneous formats.
TimeOmni-1: Multimodal time series dataset and methods for advancing complex reasoning with temporal data in large language models.
ABxLab framework evaluating LLM-powered agent decision-making in realistic consumer choice scenarios beyond task competence metrics.
NewtonBench: Benchmark for evaluating LLM agents in scientific law discovery addressing memorization, scalability, and authentic scientific process.
Systematic investigation of biases in LLM-as-a-judge systems for autonomous content evaluation with mitigation strategies.
Mechanistic analysis revealing that LLMs use compact filter head representations to encode general filtering operations in list-processing tasks.
MindPower framework enabling theory-of-mind reasoning in VLM-based embodied agents for robot-centric decision-making and action generation.
Study on using LLMs as mediators in online conflicts to foster empathy and de-escalate flame wars instead of just moderating.
STAR: Method for distilling function calling capabilities from large LLMs into smaller models using similarity-guided teacher assistance.
BrowseComp benchmark for evaluating multimodal browsing agents with visual verification and deep search capabilities in open-world environments.
First practical application of LLMs for automated microfluidic netlist generation, connecting practitioners with design automation techniques.
Theoretical framework for variable partition attribution in Shapley value-based explainable AI methods addressing attribution conflicts.
Fine-tuned DeBERTa system augmented with humor and riddle data to improve lateral thinking and creative reasoning in language models.
Investigation of LLM capabilities for code optimization and performance enhancement using problem-oriented perspective and anchor verification.
Framework for modular composition and reliable execution of LLM-enabled software with enforced contracts.
Gradient-based attribution method for explaining deep neural network decisions.
Method for improving vision-language alignment using Cauchy-Schwarz divergence.
Systematic review of NLP progress and challenges for Yoruba language.
Semantic parallelism optimization for efficient MoE model inference across multiple devices.
Diffusion-based recommendation system using continuous tokens with LLM integration.
Evaluation of LLM accuracy for health advice across languages and contextual factors.