Proposes fine-tuning and few-shot prompting approach for reference-free financial misinformation detection using LLMs without external evidence.
Presents AIPC, an AI agent-driven approach automating edge model deployment workflows including conversion, quantization, and runtime integration.
Proposes bounded-autonomy architecture using typed action contracts for safe LLM-based enterprise software interfaces, preventing model errors.
Introduces DyMETER framework for online anomaly detection adapting to concept drift in data streams without costly retraining.
Proposes semantic parsing approach for Knowledge Graph Question Answering handling negative constraints, addressing LLM hallucination limitations.
Presents Paza, a zero-shot retail theft detection system orchestrating multiple vision models without training, achieving cost-effective concealment detection.
Introduces ClimateCause dataset of complex causal structures from climate reports, including implicit and nested causality for reasoning tasks.
Studies how schema wording influences LLM behavior in structured generation under constrained decoding for JSON/XML output formats.
Presents MetaDent, a large-scale annotated dental image dataset with vision-language model benchmarks for intraoral photography analysis.
Studies feedback-based automated verification of LLM-generated code in collective adaptive systems without human code inspection, examining reliability challenges.
Proposes GenRec, a generative retrieval framework for large-scale recommendation systems using next-token prediction with handling for pagination and long sequences.
Introduces SOLIS framework for learning interpretable neural surrogates of nonlinear systems, balancing physics interpretability with model expressiveness.
Proposes RACER, combining retrieval-augmented and logits-based speculative decoding to reduce LLM inference latency while maintaining output quality.
Analyzes reasoning dynamics across 18 vision-language models, tracking confidence during chain-of-thought reasoning and measuring visual vs textual information integration.
Evaluates LLMs as adjudicators for medical diagnosis scoring on 3333 real-world hospital cases, comparing performance against expert clinician panels.
Proposes Rejection-Gated Policy Optimization (RGPO), replacing importance sampling ratios with differentiable acceptance gates for policy optimization in reinforcement learning.
Improves sparse autoencoders for foundation model interpretability using dynamic attention to optimize sparsity levels per neuron.
Presents RaTA-Tool, a retrieval-based method for tool selection in multimodal LLMs to improve complex task solving through external resource invocation.
Augments Disjoint LinUCB contextual bandit algorithm with LLM-generated pseudo-observations and calibration gates to address cold-start regret.
Proposes UniDoc-RL, a reinforcement learning framework for visual RAG using hierarchical actions and dense rewards to improve fine-grained visual reasoning.
Examines explainability challenges in scaling agentic AI adoption, addressing governance gaps and 'Agent Sprawl' phenomenon in enterprise deployments.
Attack method using adversarial suffix optimization to manipulate cost-aware LLM routers into selecting expensive high-capability models.
Study analyzing disagreement among fairness metrics for evaluating bias in machine learning systems across different demographic groups.
Framework and tools for multi-agent experiments with humans and AI agents, reducing barriers for researchers studying social decision-making.
Verifiable gradient inversion attack in federated learning that reconstructs training samples from shared gradients with certification methods.
Self-evolving logic synthesis framework using LLM agents to autonomously improve ABC codebase for circuit design.
Method for uncertainty quantification in long-form LLM generation via interrogative approach to assess semantic coherence and factuality.
Study showing RLVR-trained LLMs learn to game verifiers instead of generalizable patterns, exhibiting reward hacking behavior on reasoning tasks.
K-Token Merging technique for compressing long prompts in LLMs by reducing tokens in latent embedding space rather than token space.
Method for machine unlearning in neural networks via removing forget-specific representational directions rather than classifier suppression.
Single-layer Mamba framework for time series classification with minimal architecture redesign and TSC-specific optimizations.
System for serving complex agentic workflows with multiple LLMs and tools, handling unpredictable execution patterns and GPU constraints.
Framework for optimizing visual token pruning configurations in vision-language models using Pareto-frontier learning.
Empirical evaluation of AI tools for requirements engineering tasks against expert judgment using INCOSE criteria.
Methodological proposal for agentic AI safety research at population level, addressing risks from multi-agent interaction and collective behavior.
Benchmark study of LLM agent cooperation in social dilemmas, showing stronger reasoning models defect more in prisoner's dilemma and public goods games.
SegWithU framework for uncertainty estimation in medical image segmentation using perturbation energy in single-forward-pass inference.
Prism symbolic superoptimizer for tensor programs using hierarchical sGraph representation to optimize families of implementations.
Analysis of why vision-language models underperform on human emotion recognition compared to specialized vision classifiers.
AD4AD benchmark for evaluating visual anomaly detection models in autonomous driving under distribution shift and edge cases.
MM-WebAgent hierarchical multimodal agent for automated webpage generation maintaining style consistency across AIGC-generated elements.
Deep learning approach for constructing stochastic local search SAT solvers with performance bounds on NP-complete problems.
Deep Q-Learning implementation for autonomous driving in 2D simulated environment with custom Pygame-based track.
Study showing slower reasoning in multimodal models doesn't necessarily improve truthfulness, demonstrating inverse scaling in reasoning paradigms.
KnowRL framework combining reinforcement learning with knowledge supervision to reduce hallucination in LLMs by improving factual reasoning.
Multi-agent reinforcement learning approach for order dispatch optimization in ride-sharing systems using one-step policy optimization.
NaturalGAIA benchmark for evaluating LLM-driven GUI agents on long-horizon tasks with verifiable real-world human interaction intents.
NEMORI framework for adaptive memory distillation in LLM agents that learns what information to retain based on future utility rather than predefined heuristics.
MetaMuse uses LLMs to generate novel system algorithms through creative ideation overcoming bias toward generic heuristic designs.
DAG-based approach for automatically discovering meta reasoning skeletons to guide LLM reasoning adapted to query-specific requirements.