[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
Analysis of self-supervised speech models across 96 languages showing phonological vector arithmetic structure.
Analysis of self-supervised speech models across 96 languages showing phonological vector arithmetic structure.
Case study using ChatGPT-5 thinking mode to resolve mathematical conjecture on spectral regions.
Analysis of why LLM caching methods fail and proposes intent canonicalization with few-shot learning for cost reduction.
Convergence analysis of Stochastic Mirror Descent with matrix parameters in overparameterized regime.
Analysis of language agent failures on tool-use tasks caused by canonical path deviation rather than capability limitations.
IAPO: information-theoretic post-training framework for optimizing token efficiency in LLM reasoning chains.
Event-triggered gossip framework for distributed learning that reduces inter-node communication overhead.
Mechanistic interpretability study of how LLMs represent and use valence-related information in decision-making tasks.
Benchmark for Multi-Agent Reinforcement Learning algorithms on urban energy management tasks using CityLearn environment.
Mechanistic analysis of procedural hallucinations in LLMs showing attention and readout-stage routing errors cause value retrieval failures.
Theoretical study of how low-precision training affects scaling laws in high-dimensional linear regression with implications for quantization.
Uses soft mixture-of-experts RL for exploration in directed controller synthesis with improved zero-shot generalization to larger systems.
Bayesian nonparametric approach for predictive maintenance handling unknown failure modes in manufacturing without labeled data.
TOPReward uses token probabilities from Vision-Language-Action models as zero-shot reward signals for robotics reinforcement learning without task-specific training.
US-JEPA adapts Joint-Embedding Predictive Architectures for medical ultrasound by predicting masked latent representations instead of raw pixels.
SplitLight open-source toolkit addresses reproducibility issues in recommender systems evaluation by exploring dataset preparation and splitting strategies.
MentalBlackboard benchmark evaluates Vision-Language Models on spatial visualization tasks like paper folding and hole punching.
Value-guided multi-path reflection method for optimizing Vision-Language Models in complex robotic manipulation through improved planning and reasoning.
Multi-armed bandit approach for adaptive data augmentation to improve implicit pattern recognition in vision and language models.
IR³ framework detects and mitigates reward hacking in RLHF by reverse-engineering and interpreting internalized objectives in LLM alignment.
OptiRepair uses LLM agents to diagnose and repair infeasible supply chain optimization models through closed-loop diagnosis and domain-agnostic repair.
Human-guided agentic AI system for multimodal clinical prediction, combining autonomous workflows with domain expertise on AgentDS Healthcare benchmark.
Hierarchical Mixture-of-Agents architecture using lightweight router for cost-optimized LLM inference trading off accuracy and computational expense.
Adaptive Rejection Sampling framework for selective reasoning in LLMs, reducing token waste on simple requests while maintaining chain-of-thought benefits.
Cost-aware active search algorithm for autonomous agents to balance exploration and exploitation with unknown target recovery.
Investigates HTML-to-text extraction methods for LLM pretraining datasets, showing single extractors lead to suboptimal web data coverage.
Design principles for integrating LLMs into automotive system engineering with focus on trustworthiness and verification in safety-critical pipelines.
Proposes denoising particle filters for robot state estimation trained on single-step objectives rather than sequence unrolling.
Federated learning approach for personalized longitudinal medical report generation respecting privacy with temporal dynamics.
SkillOrchestra system for routing across multiple AI agents via learned skill transfer, reducing routing collapse in multi-turn conversations.
Vision-Language-Action models with universal pose pretraining for improved 3D state perception and embodied AI tasks.
Enables gradient-based learning for continuous-time Markov chains by decoupling simulation from differentiation for optimization.
Survey on meta-learning and meta-reinforcement learning techniques for rapid adaptation to novel tasks with minimal data.
Uplift learning method for estimating causal effects of combinatorial treatments using permutation-invariant representations.
arXiv paper proposing gradient-based severity labeling for contrastive learning in medical image biomarker classification.
arXiv paper on private inference protocols resilient to malicious client behavior in ML deployment.
arXiv paper on event-driven trading using LLMs and hierarchical-gated reward modeling with textual market signals.
arXiv paper proposing lifelong adaptability in imitation learning to improve compositional generalization beyond memorization.
arXiv paper on normal behavior modeling for time-series forecasting of telescope monitoring data.
arXiv paper on feature importance estimation and selection biases in large-scale recommender systems.
arXiv paper addressing modality gap in CLIP-based multimodal learning for medical image representation alignment.
arXiv paper on out-of-distribution detection in deep neural networks and how regional focus affects OOD detection performance.
Research on descent-guided policy gradients for multi-agent reinforcement learning addressing cross-agent noise scaling issues.
Human-AI collaboration framework using adaptive ensembles that align with human judgment when beneficial and complement when needed.
Autonomous scaling method for synthetic training environments using verifier-based reinforcement learning for reasoning language models.
Investigation of LLM knowledge encoding using nanochat family with fully open pre-training data to understand where parametric knowledge originates.
Post-calibration uncertainty quantification method balancing aleatoric and epistemic uncertainty in classifier ensembles.
Security vulnerability analysis of LLM agents exploiting skill files for prompt injection attacks through the agent skills supply chain.
Benchmark suite for evaluating video reasoning capabilities of models, focusing on spatiotemporal consistency and reasoning beyond visual quality.
Investigation of loss surface sharpness relationship to neural representation compression through local volumetric ratio and activation concentration measures.