Vision-language model robustness improvement through multimodal interaction analysis to reduce hallucinations and handle corrupted modalities.
Theoretical analysis of spurious correlation learning in preference optimization methods like DPO, characterizing mechanisms, consequences, and mitigation strategies for LLM alignment.
Characterizes geospatial web search queries at scale using embeddings and clustering, showing place-related queries exceed traditional GIS labeling schemes.
Theoretical analysis of selection bias in observational studies and methods for recovering causal effects from biased sub-populations.
MAVEN uses multi-agent prompt refinement framework with specialized agents handling person, action, and location dimensions to improve cultural fidelity in text-to-video generation.
SEMA-RAG extends RAG with self-evolving multi-agent framework for medical reasoning, enabling multi-round retrieval aligned with clinical decision processes.
FML-bench isolates AI research agent strategy impact by separating search topology from execution infrastructure, evaluating hill-climbing vs tree search.
IBAL framework robustifies multi-agent RL by addressing interaction-breaking attacks that corrupt coordination structures between agents.
PROWL improves world model robustness by actively eliciting failures on rare, interaction-critical transitions using regret-driven optimization.
Proposes block-based double decoders architecture combining decoder-only and encoder-decoder benefits with full loss supervision and efficient sequence packing.
Compares chunking strategies for RAG on German legal code, benchmarking structural, semantic, and hierarchical retrieval approaches.
Proposes importance smoothing technique for efficiently training deep state space models at scale using variational optimization.
Agent JIT compilation optimizes computer-use agent latency by compiling multi-step action sequences into efficient execution plans, reducing LLM calls.
Analyzes model distillation as minimax game between utility-constrained teacher and adaptive student, proposing tractable defense strategies.
Studies covert dialect bias in language models, showing side-by-side comparisons amplify disparities in how LMs associate traits with dialectal variations.
Proposes ReWA algorithm for sparse optimization using reparameterization, weight decay, and adaptive learning rates to address gradient unboundedness.
Reframes efficient LLM benchmarking as multiple regression with feature selection, improving prediction of full benchmark scores from partial question subsets.
GEM reformulates LLM pre-training data curation as variational optimization on hypersphere to address ontological misalignment and embedding anisotropy in data mixing.
Pair-In, Pair-Out proposes latent multi-token prediction combining input-side compression and output-side efficiency to reduce LLM inference costs.
NRLB is a multi-agent framework for plain language summarization that adapts outputs for diverse reader groups with different linguistic and cognitive abilities.
Head-to-head comparison of Claude Code and Codex executing autonomous gravitational wave data analysis pipelines on shared infrastructure without human intervention.
SafeRx-Agent proposes a multi-agent framework combining LLMs with safety verification for medication recommendation, addressing explainability and traceability in clinical decision-making.
Studies compute allocation strategies in LLM-guided evolutionary search across depth-breadth tradeoff using multi-armed bandit approach.
Pocket-Dentist: efficient multimodal LLM for on-device dental image analysis optimizing for inference speed and privacy.
Efficient sparse coding approach for multi-vector retrieval replacing k-means clustering with single-stage processing.
Neural network verification method using partial multi-neuron relaxation to formally guarantee safety properties of DNNs.
QASM-Eval dataset trains LLMs on OpenQASM-3 quantum programming including error correction, timing, and pulse-level operations.
Deep learning benchmark for predicting hip muscle forces from gait kinematics across cadence variations in 60 subjects.
Unicorn framework for scalable multi-dataset time series forecasting balancing channel-independence and channel-dependency modeling.
Multi-model study of linear representations in LLM deceptive alignment using synthetic dishonesty as controlled testbed.
Proposes RBF networks as alternative LLM architecture without deep neural networks claiming improved explainability.
Generative model for synthesizing fMRI brain imaging time series using wavelet transforms and spectral flow matching.
Maritime anomaly detection framework for AIS vessel behavior using unsupervised learning with new evaluation metric MADQI.
NumLeak measurement framework detecting memorized benchmark data in frontier LLMs via API probes and white-box analysis.
LongDS-Bench evaluates long-horizon agentic data analysis with 68 real-world Kaggle-based tasks testing context tracking across multi-turn interactions.
ML research extending calibration theory to probabilistic label ranking prediction tasks with structured output spaces.
Framework for evaluating LLM distillation beyond output matching using bounded behavioral indistinguishability formal measure.
VeriGate improves GRPO reasoning model training by adding step-level verifier feedback to address sparse supervision and credit assignment.
Unified theoretical framework analyzing gradient aggregation methods in multi-objective optimization with convergence rate analysis.
Neuro-symbolic learning framework integrating neural networks with differentiable optimization for incorporating domain knowledge as logical rules.
Distributed multi-agent reinforcement learning approach for constrained coordination problems with separable agent dynamics.
ML research on identifying datasets through semantic correlation traces in trained models via membership inference techniques.
Model extraction attack against graph neural networks using explainability interfaces in GMLaaS platforms.
Theoretical framework for universal multiclass transductive online learning with unbounded label spaces.
Mechanistic interpretability study recovering explicit Zeta map algorithm on Dyck paths from trained transformer.
Graph-conditioned mixture of experts framework for spatio-temporal traffic forecasting with node-wise expert specialization.
Balanced benchmark and unlearning method for evaluating machine unlearning across causal and relational knowledge.
Study of representation collapse during sequential post-training of LLMs with measurement suite for hidden states and LoRA updates.
Analysis of AI-style alignment signatures in LLMs through measurement and localization of post-training effects on representations.
Study of long-term effects of data selection strategies in multi-stage LLM fine-tuning and model adaptability.