Hidden States as Early Signals: Step-level Trace Evaluation and Pruning for Efficient Test-Time Scaling
Hidden States Early Signal method using step-level trace pruning to reduce computation in LLM test-time scaling through reasoning steps.
Hidden States Early Signal method using step-level trace pruning to reduce computation in LLM test-time scaling through reasoning steps.
Physics-Guided Tiny-Mamba Transformer for early fault detection in rotating machinery under domain shift and class imbalance.
Decomposition-based causal discovery method for non-stationary temporal data with autocorrelation in finance, climate, and healthcare.
SnapMLA: Hardware-optimized FP8 quantized pipelining for efficient DeepSeek MLA decoding with long context, addressing numerical heterogeneity challenges.
Accelerated prompt stress testing methodology evaluating LLM safety under repeated identical/similar prompts, identifying operational failure modes.
Framework using variability modeling to optimize LLM inference hyperparameters for energy efficiency and computational cost reduction.
ECO: Efficient neural combinatorial optimization combining batched preference optimization with Mamba, decoupling trajectory generation from gradient updates.
RDB-PFN: First relational database foundation model trained on synthetic data, enabling in-context learning for private heterogeneous databases.
Cornserve: Distributed serving system for any-to-any multimodal models handling variable input/output modalities with different computational paths.
MobileLLM-Flash: Hardware-in-the-loop architecture search for designing efficient on-device LLMs with real-time latency constraints for mobile deployment.
Explains pattern formation in diffusion models as out-of-equilibrium phase transitions driven by denoising dynamics instabilities.
Theoretical framework explaining how diffusion models learn low-dimensional manifold data through score decomposition and geometric analysis.
Analyzes capacity scaling advantages of spectral optimizers like Muon in language model training through linear associative memory theoretical framework.
Shows online recurrent learning can achieve full RTRL performance without Jacobian propagation, reducing memory from O(n^4) to O(n^2) per step.
Formalizes the Rashomon set for dimension reduction, showing multiple equally valid embeddings lead to more robust representations.
Policy Improvement Reinforcement Learning method verifying actual performance gains in LLM post-training, replacing open-loop optimization with closed-loop verification.
Statistical physics analysis of random feature models, studying training error and generalization beyond mean kernel approximation theory.
JumpLoRA framework for efficient continual learning in LLMs using sparse low-rank adapters with adaptive interference mitigation to prevent catastrophic forgetting.
LLM-based automated optimization of circuit performance, power, and area using contrastive learning of optimization rules for RTL design.
Analysis of sycophancy in LLMs showing models detect false statements but agree anyway, identifying shared circuits across twelve models responsible for this behavior.
Framework for constructing new RL policies from library of pre-trained policies offline without environment interaction under support constraints.
Semi-supervised learning approach extending FixMatch with geometric representation shaping for image classification with limited labeled data.
Hybrid technique embedding autoregressive transformers in finite element schemes for stable long-horizon forecasting of chaotic dynamical systems.
Studies adversarial attacks on LLMs through poisoned pretraining data distributed across stealth websites, examining vulnerability surface during training.
Analysis and mitigation of self-preference bias in LLM evaluators. Research on improving reliability of LLM-as-Judge evaluation systems.
Formal theorem proving for optimization problems using continual training. Developer tool for mathematical reasoning and formal verification.
Temporal curriculum for on-policy distillation in multi-turn AI agents. Research on improving agent learning through knowledge distillation.
Mathematical framework for understanding emergent intelligence and scaling laws in foundation models. Theoretical foundation model research connecting to scaling behavior.
LLM-based molecular discovery with fine-grained alignment between molecule structures and text descriptions. LLM application for chemistry/materials science.
RL-based scenario generation for testing autonomous vehicle requirements. RL application for AV testing with multi-objective trade-offs.
Deep RL with xLSTM networks for automated stock trading. RL research application with limited relevance to AI agents or LLM tools.
Efficient text-to-SQL generation without CoT or fine-tuning, reducing inference cost and complexity. Practical LLM application for code generation.
Statistical framework for detecting hallucinations in LLMs using multiple testing. Research addressing fundamental LLM reliability challenge.
Thought templates for improving reasoning in long-context LMs on multi-hop tasks. Novel technique for enhancing LLM reasoning with retrieved documents.
Structured pruning methods for LLMs focusing on task-specific optimization over layer-wise approaches. Practical technique for efficient LLM deployment.
Voyager: training-free method to generate diverse synthetic datasets using LLMs by optimizing mathematical diversity measures.
Benchmark for audio question answering that includes unanswerable questions to test LLM reliability on audio understanding tasks.
Theoretical framework proving multi-layer cross-attention optimality for multi-modal in-context learning in transformers.
Inter-Layer Structural Encoders framework aggregating intermediate layer representations for improved LLM predictions with minimal parameters.
Comparative analysis of datasets, foundation models, and gaps preventing medical AGI, focusing on surgical AI benchmarks.
Score calibration method using percentile-rank normalization to fuse vector and graph-based retrieval signals for multi-hop QA.
Analysis of modality gap in multi-modal models like CLIP from robustness perspective, examining whether gap should be closed.
Vision-Language-Action model decoupling intent and action via latent world modeling to improve VLM utilization in robotics.
Sequence modeling approach using complex Hilbert space and quantum-inspired mechanisms for contextual semantic meaning in NLP.
Two-hop QA retrieval framework formalizing regime-conditional splitting and transferable router for question-dominant vs bridge-dominant queries.
Training method unifying supervised fine-tuning and RL for LLMs via group advantage estimation and dynamic coefficient rectification.
Multi-armed bandit algorithm optimizing concave statistical utility functions using influence-function gradients for non-expected-reward objectives.
Audio2Tool dataset with 30k queries benchmarking speech-based tool-calling performance in Speech Language Models across diverse domains.
Mixture-of-Experts architecture with heterogeneous expert sizes that adapts computational costs to varying token-level complexity in LLMs.
Startup reports $67k Gemini API charge accrued in 19 hours following API changes, with billing dispute escalation.