VCoT-Bench evaluates LLMs on Rust program verification via chain-of-thought reasoning, testing logical deduction abilities beyond binary pass/fail.
Method for reliable uncertainty quantification in Vision-Language-Action models by shifting focus to safety-critical moments in robotic control.
PowerFlow applies principled distribution matching to unsupervised reinforcement learning from LLM internal feedback without external supervision.
TARo enables frozen LLMs to perform structured reasoning at inference time through token-level adaptive routing, avoiding expensive post-training alignment.
Unsupervised discovery of transition-structure concepts in text via temporal co-occurrence patterns using contrastive learning on large corpus.
Adaptive context allocation method for LLM long-context inference using uncertainty-triggered token-level budgeting to address attention dilution.
Vision-language model method for temporal out-of-distribution detection and domain generalization in open-world settings using adaptive pattern matching.
Analysis of how standard LLM decoding strategies (top-k, nucleus sampling) exclude contextually appropriate but statistically rare tokens compared to human language production.
ICE-Guard detects spurious feature reliance in LLM decision-making through intervention consistency testing on demographic, authority, and framing features.
Addresses sim-to-real transfer for vision-language-action models in robotics by generating diverse 3D simulation worlds for RL fine-tuning.
iSatCR optimizes onboard computing and routing for LEO satellite data processing using graph neural networks to reduce ground transmission bottlenecks.
CausalVAD applies causal intervention to de-confound end-to-end autonomous driving models, addressing dataset bias and improving reliability.
ICE framework evaluates explanation faithfulness in LLMs via randomization tests with multiple intervention operators, distinguishing genuine faithfulness from chance.
Memento-Skills introduces an LLM agent that autonomously designs and improves task-specific agents through continual learning with stateful prompts and reusable skills.
Proposes variational guidance for autonomous aerial vehicle trajectory learning to address credit assignment and training instability in sparse reward RL settings.
BeamAgent combines LLMs with wireless beamforming optimization through decoupled intent parsing and alternating optimization, separating LLM reasoning from numerical computation.
RewardFlow proposes topology-aware reward propagation on state graphs for RL-enhanced LLM agents, addressing sparse reward limitations without expensive dedicated reward models.
Proposes evaluation framework beyond accuracy for human-AI collaborative decision-making, addressing miscalibrated reliance and team effectiveness.
Studies entropy trajectory shape in chain-of-thought reasoning to predict LLM correctness without additional inference, testing on GSM8K with Qwen2.5-7B.
Proposes unified taxonomy with 11 dimensions for categorizing deep learning approaches to multivariate time series anomaly detection.
CRAFT method for aligning diffusion models through fine-tuning, addressing limitations of SFT and DPO-style preference optimization approaches.
Hypothesis-Conditioned Query Rewriting improves RAG systems by rewriting queries to prioritize decision-relevant evidence over topical relevance.
Lightweight cryptographic framework for verifiable AI inference enabling clients to verify model outputs without rerunning computation.
SEM method for post-hoc debiasing of CLIP via sparse embedding modulation to remove social and spurious biases.
Neural network approach to autoregressive time series estimation using backpropagation while preserving interpretability.
SAVeS framework steers safety judgments in Vision-Language Models through semantic cues and textual/visual interventions.
FedTrident defends federated learning-based road classification against poisoning attacks from malicious participants.
Studies how uncertainty estimation scales with sampling in reasoning language models using self-consistency and verbalized confidence.
D5P4 framework applies determinantal point processes to discrete diffusion decoding for diverse parallel text generation.
Method for splitting pretrained language models into specialized domain-specific models using continued pretraining strategies.
Multi-agent framework for grounding vision-language navigation using probabilistic reasoning about spatial relations and metric constraints.
Evaluates State Space Models as vision encoders for Vision-Language Models, comparing SSM backbones to transformer-based alternatives.
DreamPartGen generates semantically grounded 3D objects with part-level decomposition using text-to-3D diffusion methods.
DriveTok proposes efficient 3D tokenization for multi-view driving scenes to improve autonomous driving systems and world models.
Nemotron-Cascade 2: 30B open-weight MoE LLM with strong reasoning and agentic capabilities, achieving IMO Gold Medal performance.
NavTrust benchmark evaluates trustworthiness of embodied navigation agents under real-world corruptions in Vision-Language Navigation and Object-Goal Navigation tasks.
Establishes improved learning rates for stochastic gradient descent and Nesterov accelerated gradient with generalization performance guarantees.
Chat Incremental Pattern Constructor extracts ordered token-transition rules from text for interpretable machine learning rule extraction.
Optimization methods for inverse classification problems including counterfactual explanations and adversarial examples using logistic and softmax classifiers.
CADGL uses context-aware deep graph learning for predicting drug-drug interactions with improved generalization and robustness.
μLO derives Maximal Update Parametrization for learned optimizers to improve meta-generalization across network widths and unseen tasks.
Flow matching approach with large-scale synthetic dataset for solving inverse ellipsometry problem of reconstructing optical film properties.
ODE-constrained generative model for synthesizing realistic 12-lead ECG training data to address scarcity of labeled medical recordings.
Cliqueformer uses structured transformers for model-based optimization in design problems like protein engineering via offline learning.
VOGP algorithm using Gaussian process bandits for black-box vector optimization with incomplete order relations and Pareto optimality guarantees.
Theoretical analysis showing shallow nonlinear networks learn linearly separable features with polynomial width scaling relative to data dimension.
Methods to achieve real-world efficiency gains from token filtering in LLM training through improved sparsity and adaptive filtering strategies.
Survey of Part-Prototype Models for explainable AI, examining interpretability mechanisms and competitive limitations versus alternative approaches.
Two neural architectures for precipitation nowcasting integrating weather station data and radar measurements for improved forecast skill.
OPUS-VFL addresses privacy-utility tradeoffs and incentive mechanisms in Vertical Federated Learning with heterogeneous client resources.