Optimal FALQON for Quantum Approximate Optimization via Layer-wise Parameter Tuning
Optimal FALQON improves convergence speed for quantum approximate optimization on NISQ devices through layer-wise parameter tuning.
Optimal FALQON improves convergence speed for quantum approximate optimization on NISQ devices through layer-wise parameter tuning.
CDS4RAG framework for cyclic dual-sequential optimization of both retriever and generator hyperparameters in Retrieval-Augmented Generation systems.
Methodology for detecting hallucinations in LLM chain-of-thought reasoning by distinguishing actual reasoning evaluation from answer-surface correlates.
Research on reinforcement learning challenges for scalable and trustworthy intelligent systems in distributed environments and LLM post-training.
Empirical study examining software engineering discourse produced by autonomous AI agents interacting primarily with each other on MoltBook network.
Zero-shot composed image retrieval method that decouples endpoint and semantic transition learning for lightweight projection-based approaches.
AIPO method for improving LLM reasoning through active interaction and reinforcement learning, extending beyond policy model capability boundaries.
Study of using multimodal LLMs with remote sensing imagery for smart city tasks including design suggestions and risk identification.
Research proving mechanism design alone is insufficient for safe AI agent cooperation, proposing prosocial agent approaches for beneficial multi-agent interaction.
Framework for evaluating LLM calibration in open-ended question answering tasks, addressing reliable deployment in high-stakes domains.
MEEC-Net uses meshfree exterior calculus to learn physics from point clouds with data efficiency and transferability across resolutions and parameters.
Magis-Bench benchmark for evaluating LLMs on magistrate-level legal judgment tasks including weighing claims and rendering reasoned decisions.
DUET framework for optimizing token-budget allocation in reinforcement learning with verifiable rewards by jointly tuning rollout distribution across prompts.
Systematic evaluation of defense mechanisms against persistent memory attacks on stateful LLM agents across architectural layers and multiple open-source models.
Research on zero-shot imitation learning for training agents from fixed demonstration datasets to solve new tasks without complete task-specific examples.
Analyzes sinks and diagonal patterns in attention mechanisms as methods for preventing oversmoothing in deep networks.
Proposes method for recovering continuous-time dynamics from discrete observations using semi-group property constraint instead of local supervision.
Studies security risks of multi-agent LLM networks where agents spawn subagents, modeling attack vectors in hierarchical delegation.
Evaluates whether LLM benchmarks underestimate performance by comparing human annotations with LLM predictions on hallucination detection in RAG.
Presents system for frozen-LLM coding agents with validation-grounded memory and adaptive retrieval for execution-feedback repair cycles.
Analyzes how transformers with softmax attention implement preconditioned Richardson iteration for in-context Gaussian kernel regression.
Introduces MathConstraint, adaptive benchmark for evaluating LLM combinatorial reasoning with solver-based verification and dynamic difficulty.
Proposes multi-level graph attention network with contrastive learning for knowledge-aware recommendation systems using knowledge graphs.
Analyzes softmax attention scaling limits to understand when selectivity emerges versus uniform averaging in long-context transformers.
Demonstrates safety alignment vulnerabilities in LLMs by targeting single neurons that control refusal and harmful knowledge to bypass safety measures.
Proposes MARLaaS system for concurrent multi-tenant reinforcement learning fine-tuning of LLMs using verifiable rewards for agentic capabilities.
Studies continuity as inductive bias in sequential models, analyzing whether continuous-time formulations like state-space models behave continuously.
Introduces VeriContest benchmark for evaluating LLMs on verifiable code generation with formal specifications and machine-checkable proofs.
Proposes mixed-curvature Riemannian flow matching for parameter-efficient adaptation of vision models using geometry-aware few-shot learning.
Technical report on ZAYA1-VL-8B, a compact mixture-of-experts vision-language model achieving competitive performance with smaller parameter count.
Explores intra-expert activation sparsity in Mixture-of-Experts LLMs to improve computational efficiency beyond standard sparse expert activation.
Analyzes scaling behavior of minimalist transformer world models on Atari 100k benchmark to understand data efficiency in generalist systems.
Addresses context compression validation for long-horizon LLM agents by analyzing accuracy gaps when trajectory summaries are generated asynchronously.
Applies differentiable prompt tuning and in-context learning to medical imaging for Alzheimer's diagnosis using multimodal models.
Evaluates epistemic overreach in LLM explanations of sensor data, measuring when generated accounts exceed available evidence in personal sensing applications.
Lattice Deduction Transformer uses recurrent attention with lattice projections for logical reasoning tasks, achieving perfect accuracy on constraint satisfaction with minimal parameters.
STRIDE combines time series foundation models with LLM reasoning to improve forecasting interpretability while handling continuous numerical values without excessive tokenization.
Introduces PARD-2, a draft model optimization for speculative decoding that aligns training with inference goal of maximizing consecutive token acceptance in LLM inference.
Proposes KeyStone, an inference-time self-consistency method using geometry guidance for physical AI models that generate action trajectories via diffusion/flow matching.
Introduces AgentCollabBench, a diagnostic benchmark for measuring multi-hop process failures in multi-agent systems where individual agents appear correct but collaboration silently fails.
Proposes SketchVerify for structured inference-time scaling where LLMs enumerate algorithmic strategies and write program sketches verified by execution for cost-effective improvement.
Empirically analyzes fairness in LLM explanations across demographic groups, introducing Explanation Fairness Taxonomy framework for evaluating explanation quality disparities.
Survey of attention-based graph neural networks covering mechanisms for adaptive feature selection and noise filtering in graph representation learning.
Compares voting strategies for LLM code generation including Semantic Voting, analyzing how textual voting, ranking, and execution-based agreement contribute to code selection.
Introduces PrepBench benchmark evaluating natural language-driven data preparation using LLMs, assessing progress toward replacing GUI-based data transformation tools.
Introduces Gate-and-Merge, a zero-shot framework for compositional personalization of vision-language models using lightweight LoRA adapters paired with concept tokens.
Develops end-to-end autonomous parking using reinforcement learning with Gaussian Splatting simulator for real-to-sim-to-real transfer in extreme parking scenarios.
Proposes AgentForesight for early failure prediction in LLM-based multi-agent systems through online auditing, enabling intervention before trajectory-level failures cascade.
Introduces PROBE, a framework for structured recovery in software engineering agents that converts runtime failures into grounded recovery guidance for subsequent attempts.
Proposes AdaPreLoRA, an optimization method for Low-Rank Adaptation using Adafactor preconditioning to handle rank-deficiency issues in LoRA factor-space preconditioners.