Theoretical analysis of classifier-free guidance in diffusion models, studying why inference-time guidance is necessary for generative modeling.
Gradient-based data selection framework for online LLM fine-tuning that accounts for sequential data arrival and adaptive optimizers.
AgenticFlict: Large-scale dataset of 1,000+ merge conflicts from AI coding agent pull requests on GitHub, analyzing integration challenges as agents become active contributors.
AMuFC framework for multimodal fact-checking that adaptively determines when visual evidence improves or reduces accuracy, challenging assumption of universal multimodal benefit.
LINE: Training-free iterative method using LLMs to explain individual neurons in vision models, improving interpretability beyond predefined concept vocabularies.
Vision-language models enhanced with skill conditioning for geographic reasoning and self-evolution in image geolocation tasks, addressing hallucination and outdated knowledge issues.
Mosaic: probabilistic weather forecasting model addressing spectral degradation and aliasing in ML-based weather prediction.
AmaraSpatial-10K: 10,000+ metric-scaled 3D assets optimized for zero-shot deployment in embodied AI and spatial computing applications.
Evaluation of prompt injection defenses across 20,000 adaptive attacks on nine defense configurations; only output filtering remained unbroken.
AgenticRecTune: multi-agent framework with self-evolving skill hub for optimizing multi-stage recommendation system pipelines.
COHERENCE benchmark for evaluating multimodal LLMs on fine-grained image-text alignment in interleaved document-like contexts.
Multi-agent LLM benchmark for negotiation that tests dynamic grounding and communication repair across conversational turns.
Gyan: neuro-symbolic language model combining transformers with symbolic reasoning to improve compositionality, interpretability, and reduce hallucinations.
Framework for evaluating agentic stock prediction systems using LLM judges and closed-loop reinforcement learning feedback on behavioral dimensions.
Asymmetric on-policy distillation method for training student language models with token-level teacher feedback, improving upon standard RL and off-policy approaches.
AI CFD Scientist: open-source physics-aware AI agent for autonomous computational fluid dynamics discovery with LLM-based scientific reasoning loop.
Benchmark for LLM-assisted formal mathematical reasoning using Lean and Mathlib, evaluating pull request merge-readiness for library contributions.
Study on when neural networks fail to extrapolate out-of-distribution, analyzing feature learning versus data-generating-process identifiability.
Mechanistic analysis of hallucination failures in vision-language models, tracing issues to geometric over-alignment in decoder-based VLMs.
FactoryNet: industrial time-series pretraining dataset with 51M datapoints across 23k task executions for zero-shot transfer and anomaly detection.
Research on FP4 quantization for pretraining large language models, investigating stability and convergence issues in full-pipeline low-precision training on Llama 3.1-8B.
Key-Value Means presents block-recurrent attention mechanism achieving O(N) complexity with fixed or growing state for efficient long-context transformer inference.
Tool combining weakest-precondition analysis with agentic Claude Code CLI for specification inference in Move Prover to reduce verification boilerplate.
Metis framework reformulates LLM jailbreaking as inference-time policy optimization using adversarial POMDP for improved red teaming scalability.
CoWorld-VLA multi-expert world model framework for vision-language-action autonomous driving with planning-oriented spatiotemporal representations.
ALAM latent action model for vision-language-action models that extracts action priors from video using algebraically consistent representations.
Formal framework for probabilistic safety shielding in Markov decision processes with conservative guarantees.
DataMaster autonomous system for data-centric ML research automating dataset discovery, adaptation, validation, and knowledge propagation.
TMPO trajectory matching policy optimization for diffusion alignment addressing reward hacking through probability distribution constraints.
MCPShield attack detection framework for LLM agent tool-call traffic via Model Context Protocol, using graph encoding and embeddings.
HEPA self-supervised architecture for event prediction in multivariate time series using horizon-conditioned JEPA pretraining.
LiBaGS lightweight method for selecting informative synthetic training data by scoring boundary proximity, uncertainty, and data density.
ZeNO gradient-free noise optimization method for reward alignment in diffusion and flow models without backpropagation.
Reasoning-prefix masking technique for distilling visual-reasoning capabilities from large VLMs into compact student models.
FAMeX algorithm for AI explainability using feature association maps based on graph-theoretic approaches.
Safe reinforcement learning approach that learns when agents should act via communication-efficient timing decisions under Lyapunov safety constraints.
CAWI initialization method for randomized neural networks using copulas to capture inter-feature dependencies.
Federated multimodal graph learning approach addressing modality heterogeneity and incomplete data across distributed networks.
OceanCBM concept bottleneck model for interpretable ocean forecasting that provides mechanistic explanations aligned with physics.
Study on decision-making with AI assistance, examining how decision-makers interpret model confidence and prediction utility in high-stakes domains.
Theoretical analysis establishing population risk bounds for Kolmogorov-Arnold Networks trained with mini-batch SGD and differential privacy.
Embedding Temporal Logic framework for runtime monitoring of perception-based autonomous systems without expensive learned abstraction modules.
Multi-rollout on-policy distillation method for LLMs that leverages peer successes and failures to provide denser token-level supervision beyond sparse verifier rewards.
FPILOT framework applies inference-time optimization via Model Predictive Control to RL trading agents for portfolio management, incorporating price forecasts at deployment.
Method for robust LLM alignment using ordinal decomposition of discrete rewards in RLHF with stochastic auto-raters for long-form QA and instruction following.
IGT-OMD: implicit gradient transport for decision-focused learning with delayed feedback in online bilevel optimization.
Method for node classification in multiplex graphs with heterophily using adaptive cross-domain learning approach.
UFO: domain-unification-free neural operator framework for learning cross-domain function space mappings.
Research on how upstream training choices affect model robustness when capabilities are retained through subsequent fine-tuning.
RSNet: open-source R package for robust network inference in high-dimensional data using resampling-based framework.