IndexRAG: Bridging Facts for Cross-Document Reasoning at Index Time
IndexRAG approach for multi-hop question answering that performs cross-document reasoning at indexing time using bridge entities.
IndexRAG approach for multi-hop question answering that performs cross-document reasoning at indexing time using bridge entities.
SF-Mamba state space model for vision addressing non-causal patch interactions with improved computational efficiency.
SlideFormer system for fine-tuning large language models on single GPU via asynchronous engine and sliding window approach.
EngGPT2-16B open Italian LLM trained on 2.5T tokens, efficient inference with performance comparable to larger models.
Multi-agent reinforcement learning approach for managing delayed channel state information in multi-satellite communication systems.
Unlearning method for one-step generative models using unbalanced optimal transport for safer image generation.
LenghuSky-8 millisecond-resolution network dataset for time series foundation models with high-frequency data.
FEAT foundation model with linear complexity for structured data in healthcare, finance, and e-commerce with improved scalability.
DanceHA multi-agent framework for document-level aspect-based sentiment analysis, extracting ACOSI tuples from documents.
CompDiff uses hierarchical compositional diffusion to generate fair medical images across demographic groups and intersections.
EmoLLM framework integrates appraisal-grounded cognitive-emotional reasoning into LLMs for contextually appropriate responses.
Analysis of human-LLM chat logs characterizing delusional spirals and negative psychological effects from extended chatbot interactions.
Manifold-Matching Autoencoders regularize autoencoders by aligning pairwise distances between latent and input spaces.
Research on classifying malicious AI agent skills using repository context to improve detection in skill marketplaces.
REFORGE reveals vulnerabilities in image generation model unlearning through multi-modal adversarial attacks in black-box settings.
BATQuant proposes outlier-resilient MXFP4 quantization via learnable block-wise optimization for deploying MLLMs and LLMs on accelerators.
Analysis of multimodal LLM-generated natural language explanations for face verification on unconstrained face images.
Omanic, a benchmark for step-wise evaluation of multi-hop reasoning in LLMs with step-level annotations for diagnosing failures.
Investigation of linguistically related language guidance for LLM translation in low-resource settings without large parallel data.
Study of emergent AI agent communities on platforms, analyzing 167k+ agents learning from each other without researcher intervention.
Kestrel, a training-free method for mitigating hallucinations in large vision-language models using grounding and self-refinement.
World action models for embodied control that eliminate test-time future imagination while maintaining action performance.
Resource-aware LLM-based agent reasoning for embodied robots using reinforcement learning to balance computation and action execution.
In-context learning improvement for vision-language models using retrieved counterfactuals for better visual reasoning.
SpecMoE mixture-of-experts foundation model for cross-species EEG decoding with spectral-temporal fusion.
Formal model for selecting statements that find common ground across diverse preferences using generative AI.
TurnWiseEval benchmark and analysis of multi-turn vs single-turn LLM capabilities with step-level evaluation.
InCoder-32B, a 32B code foundation model optimized for industrial programming tasks with hardware semantics and resource constraints.
Cross-embodiment dexterous grasping policy enabling zero-shot transfer across different robot hand morphologies without retraining.
Behavior tree planning for robot manipulation using context-aware grounding to automate controller design without extensive manual effort.
Study examining reasoning mechanisms in diffusion-based video models, challenging chain-of-frames assumptions about how reasoning emerges.
Comprehensive survey of LLM reasoning covering inference scaling, learning to reason, and agentic systems as key advancement areas.
CHARM method calibrates reward models using Chatbot Arena scores to mitigate model preference bias and reward hacking in RLHF.
IMAIA interactive maps assistant enables natural language interaction with vector maps and satellite imagery with geospatial intelligence.
Survey of LLM applications in wireless communications covering adaptation, autonomy, and intelligent system design for complex communication networks.
Multi-agent pipeline for street design generation combining image generation and infrastructure design for urban planning visualization.
Hilbert system combines informal LLM reasoning with formal theorem proving in Lean 4 for verifiable mathematical proofs.
ReasoningBank memory framework enables LLM agents to learn from interaction history and distill generalizable reasoning strategies for continuous tasks.
Zephyrus agentic framework combines weather foundation models with LLM reasoning for interactive scientific workflows in meteorology.
Study showing AI agents fail under realistic user behavior variations; proposes high-fidelity human trait simulations for robust agent testing.
PREFINE framework enables personalized story generation using simulated user critics and rubric generation without explicit user feedback.
Multi-agent debate framework using small language models for cost-efficient LLM safety evaluation, with HAJailBench benchmark for jailbreak testing.
Alignment-Aware Quantization: PTQ method for efficient LLM deployment that preserves behavioral alignment and safety properties, not just minimizing reconstruction error.
SpatialBench: benchmark measuring multimodal LLM spatial cognition across hierarchical abilities for real-world physical environment interaction.
Analysis of multi-agent path finding algorithm design trade-offs under realistic robot execution constraints for warehouse and manufacturing applications.
Stepwise Think-Critique: unified framework integrating reasoning and verification in LLMs for robust, interpretable problem-solving with intertwined evaluation.
FusionRoute: token-level collaboration method enabling multiple specialized LLMs to work together, combining domain expertise efficiency with generalization.
VisTIRA: tool integration approach addressing modality gap where VLMs underperform on visual math problems compared to text-based versions.
LogicSkills: benchmark isolating three fundamental logical reasoning skills in LLMs: formal symbolization, countermodel construction, and logical inference.
Empirical study of latent chain-of-thought in LLMs using structural causal models to analyze intermediate computation steps beyond correlation-based probes.