LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence
LAMB: LLM-based audio captioning framework bridging modality gap between audio and text using Cauchy-Schwarz divergence.
LAMB: LLM-based audio captioning framework bridging modality gap between audio and text using Cauchy-Schwarz divergence.
GeoMotionGPT: LLM framework for motion understanding using geometry-aligned discrete motion tokenization and embeddings.
Imagine-then-Plan framework for agent learning using world models and adaptive lookahead for complex task planning.
RAG-3DSG: retrieval-augmented generation approach for constructing 3D scene graphs with uncertainty estimation for robotics tasks.
Multi-agent LLM framework for generating research limitations by identifying methodological issues beyond superficial statements.
Research on decomposable inference in large models showing gradient updates are localized, reducing inference costs and complexity.
Jacobian Scopes: gradient-based methods for token-level causal attribution in LLMs to identify which prior tokens influence predictions across layers and attention heads.
VibeVoice-ASR framework for speech understanding in long-form audio using single-pass processing to handle context fragmentation and multi-speaker scenarios.
NaVIDA improves vision-language navigation agents by explicitly modeling action-grounded visual dynamics for better planning and generalization in embodied environments.
Benchmark for evaluating reasoning in baby language models trained on child-directed speech; developmentally-inspired evaluation methodology.
Expert-panel study on human detection of LLM-generated Korean text using rubric-based calibration framework for attribution.
Mixture-of-Experts approach for time-series forecasting transformers using segment-wise routing to improve scaling and temporal dynamics.
MDial framework for generating multi-dialectal dialogue data; addresses LLM performance gaps for non-standard English speakers.
Adversarial attacks against search-enabled LLM fact-checking systems; proposes DECEIVE-AFC for testing robustness of retrieval-augmented verification.
First labeled dataset of 98,380 malicious agent skills characterizing security threats in LLM-based agent registries and ecosystems.
Open-source singing voice synthesis system with zero-shot generalization and controllable generation capabilities.
System for natural language graph analytics over large property graphs using LLMs; enables querying complex heterogeneous datasets efficiently.
Token reduction method for multi-modal LLMs in autonomous driving to improve efficiency while maintaining human-vehicle interaction.
Physics-based tropical cyclone estimation using spline-parameterized KAN for efficient edge device deployment on satellite data.
Framework for on-policy supervised fine-tuning of LLMs using Distribution Discriminant Theory to improve generalization over standard SFT.
Generic object tracking method using joint-embedding predictive architecture with model adaptation and occlusion reasoning.
Research on identifying missing persona dimensions for user simulation in dialogue systems to improve validity of simulation results.
Geometric analysis of optimization dynamics in grokking; shows transformers train in low-dimensional subspaces with detailed PCA findings.
Research on early-warning signals for grokking phenomenon via loss-landscape geometry analysis across sequence learning benchmarks.
Investigation of text-to-image diffusion models' effectiveness as synthetic data generators, revealing performance regression when used for training data generation.
LESA method for accelerating diffusion models through learnable stage-aware predictors that adapt to stage-dependent dynamics with feature caching.
Survey of neural routing solvers that use deep learning to tackle vehicle routing problems by learning implicit heuristic rules from data.
Study of LLM unlearning robustness in multi-turn interactive settings, addressing safety, privacy, and legal concerns in machine unlearning.
Discrete gauge-theoretic framework for understanding superposition in LLMs using sheaf theory and local semantic charts instead of global dictionaries.
arXiv paper introducing RMBench robotic manipulation benchmark emphasizing memory-dependent tasks and policy design insights.
arXiv paper proposing Causal Hamiltonian Learning Unit as deep learning primitive for temporal dynamics addressing LSTM/Neural ODE tradeoffs.
arXiv paper analyzing Google's SynthID-Text watermarking system for LLM-generated text detection using tournament-based methods.
arXiv paper evaluating multi-vendor LLM agent teams for clinical diagnosis, testing whether diversity reduces correlated failure modes.
arXiv paper on time series forecasting with multi-dimensional exogenous integration for industrial applications like aviation.
Study showing LLM-as-a-Judge evaluation frameworks are unreliable for safety assessment, demonstrating vulnerability to distribution shifts in red-teaming.
HEARTS benchmark for evaluating LLM reasoning on health time series across multiple physiological modalities and temporal dependencies.
Symbolic ML approach for failure detection in chemical processes, emphasizing interpretability and safety over neural methods; includes ethylene oxidation case study.
AI-powered platform using LLM personas to teach deliberative democratic skills and consensus-finding through simulated discussion scenarios.
ML competition for agricultural vision focusing on data-centric approaches and model generalization under distribution shifts in real field conditions.
Evaluation framework for tabular foundation models using proper scoring rules to assess full predictive distributions, not just point estimates.
Fine-tuning method for Vision Transformers using concept guidance to reduce spurious correlations and improve robustness to distribution shifts.
Clinical feasibility study of LLM-based conversational diagnostic AI (AMIE) in real primary care workflows with safety evaluation.
Study of emotion as latent factor affecting LLM reasoning and attention mechanisms, rather than just a prediction target.
Motion forecasting for autonomous vehicles handling open-world scenarios with imperfect perception and evolving object taxonomies.
Vision-language-action model for autonomous driving using perception-planning distillation to improve visual encoding and trajectory planning stability.
Research on how LLMs handle compositional language tasks (adjective-noun relationships), comparing external performance with internal model representations.
Research comparing LLM performance in healthcare triage across evaluation formats. Shows evaluation methodology significantly affects model assessment outcomes.
KEPo: research on knowledge graph poisoning attacks against GraphRAG systems. Analyzes vulnerabilities when LLMs rely on external databases.
RoboClaw: Agentic framework unifying data collection, policy learning, and deployment for long-horizon robotic manipulation using Vision-Language-Action systems.
Controlled experiments showing language models prefer correct answers due to data compressibility structure rather than truth, using small transformers on contradictory corpora.