Large Language Models Can Help Mitigate Barren Plateaus in Quantum Neural Networks
Proposes using LLMs to help mitigate barren plateaus in quantum neural network training through adaptive parameter initialization.
Proposes using LLMs to help mitigate barren plateaus in quantum neural network training through adaptive parameter initialization.
Study evaluating emergent lifelong learning behaviors in LLMs during multi-turn interactions, proposing new evaluation benchmarks for character-like consistency.
TARAC method addresses hallucinations in vision-language models by improving temporal attention mechanisms during generation without extensive retraining.
Research on energy-efficient optimization techniques for LLM deployment, including quantization and local inference strategies to reduce carbon emissions.
PODS decouples rollout generation from policy updates in LLM RL, addressing compute asymmetry through down-sampling.
LOOPE method learns optimal patch ordering in Vision Transformer positional embeddings for improved spatial information encoding.
RL^V framework unifies LLM reasoners with verifiers using value functions for improved test-time compute scaling during reasoning.
Bayesian approach for Vision Language Models to reduce hallucinations and overconfidence in VQA through selective prediction.
TokUR enables LLMs to self-assess uncertainty at token-level for improved reasoning and response reliability in multi-step tasks.
SpatialScore: comprehensive benchmark for evaluating spatial intelligence of multimodal LLMs with data-driven and agent-based assessment approaches.
GoT-R1: reinforcement learning framework enhancing multimodal LLM reasoning for complex visual generation with precise spatial relationships and attributes.
Fine-tuning approach for LLMs to predict diverse user behaviors, addressing overfitting to frequent behaviors while capturing long-tailed behavior distribution.
World models for interactive video generation with action conditioning and autoregressive decoding to support planning and future prediction.
Framework using LLMs for few-shot code generation to create safety-critical driving scenarios in CARLA simulator for autonomous driving evaluation.
LLM-based autonomous agent for power system voltage control, using experience-driven learning to generate dispatch strategies in distribution networks.
Data Mixing Agent: LLM-based method to automatically re-weight training data domains during continual pre-training, preventing catastrophic forgetting.
PRIX: efficient end-to-end autonomous driving model planning from raw camera pixels without LiDAR, reducing model size and computational requirements.
MDM-OC: framework for scalable, reversible model composition enabling continual learning without task interference or catastrophic forgetting.
Genetic programming approach for symbolic distillation of neural networks, using teacher-student smoothness alignment to improve explainable AI model accuracy.
Protocol for reliable evaluation of low-precision retrieval systems, addressing spurious ties and variability in relevance scoring with reduced numerical precision.
Analysis of LLM use in newsmaking across 40,000+ articles using AI-text detectors, showing increased GenAI adoption in local and college media.
Proximal SFT: supervised fine-tuning method using trust-region constraints to prevent capability deterioration when adapting foundation models to new tasks.
LLM-based synthetic training reduces maritime domain model costs 261x by using LLMs as teachers for small language model training.
FS-DFM enables fast long text generation using few-step diffusion language models with parallel position generation.
StyleBench evaluates trade-offs between structured reasoning styles and efficiency/robustness in LLM inference.
Position paper analyzing measurement gaps in reinforcement learning with verifiable rewards for LLMs on structured tasks.
SecureVibeBench evaluates code generation security of LLM-powered code agents against realistic vulnerability scenarios.
Mathematical framework interpreting Transformers as discretizations of integro-differential equations.
LLM-based system for generating standards-aligned math word problems customized to student interests and ability levels.
HiPRAG uses hierarchical process rewards to improve agentic RAG efficiency, reducing over-search and under-search behaviors.
Unified framework analyzing sequence models (Transformers, SSMs, gated RNNs) through coefficient dynamics lens.
Survey of inductive reasoning in LLMs, covering particular-to-general thinking patterns and knowledge generalization capabilities.
RAGen framework for generating domain-specific question-answer pairs to adapt RAG systems to specialized applications.
Risk-sensitive abstention in bandit algorithms for high-stakes AI where errors are irreparable without expert guidance.
Multi-hop reasoning over knowledge graphs using multi-view RAG with LLMs, addressing Transformer attention specialization patterns.
SimBench provides first standardized benchmark for evaluating how faithfully LLMs simulate human behaviors across diverse tasks and metrics.
AtlasKV enables RAG systems to integrate billion-scale knowledge graphs efficiently in limited VRAM by avoiding expensive external retrieval modules.
Method to automatically extract and explain what features human feedback data encodes when training language models, addressing unpredictability in RLHF approaches.
Analysis of multilingual reasoning gaps in reasoning language models, showing deficits stem from language understanding failures in low-resource languages.
Method for interpreting LLM reasoning by resampling multiple chain-of-thought branches to measure causal influence and underlying computation.
LLM-guided decompilation framework using context to improve re-executability of decompiled binaries for security analysis.
SynthAgent: Framework for web agent adaptation using synthetic data generation with quality filtering to handle hallucinations and trajectory noise.
GroupRank: Efficient passage reranking paradigm using LLMs with groupwise ranking to balance efficiency and accuracy.
LiveCLKTBench: Benchmark pipeline for reliably measuring cross-lingual knowledge transfer in multilingual LLMs with time-sensitive queries.
Framework for process-centric evaluation of agentic software systems, analyzing execution trajectories and reasoning beyond outcome metrics.
Theoretical framework for sparse dictionary learning in neural networks, analyzing piecewise biconvexity and spurious minima in mechanistic interpretability.
WisPaper: AI agent system for academic paper discovery and organization, addressing semantic search and workflow fragmentation challenges.
VPR-AttLLM framework using LLM semantic reasoning to improve geo-localization of crowdsourced flood imagery.
Multimodal RAG system enhanced with knowledge graphs for audio-visual retrieval, extending LLM capabilities to multimodal domains.
Research on variance-aware tree policies for Monte Carlo Tree Search, improving upon UCB-based methods used in AlphaZero-style algorithms.