Unified Biomolecular Trajectory Generation via Pretrained Variational Bridge
Pretrained variational bridge for efficient molecular dynamics trajectory generation across diverse molecular systems.
Pretrained variational bridge for efficient molecular dynamics trajectory generation across diverse molecular systems.
Automated detection pipeline for unverbalized biases in LLM chain-of-thought reasoning without predefined categories.
Performance characterization framework for small language models on edge devices using Roofline model analysis.
Benchmarking framework for time series foundation models addressing data quality, task alignment, and evaluation rigor.
Framework for measuring propensities (behavioral tendencies) in AI models beyond capability assessment using Item Response Theory.
Theoretical analysis of gradient descent convergence rates under large step sizes in separable logistic regression.
Discrete diffusion framework using sample-efficient conditional probability estimators for discrete state space generation.
Analysis showing test-time training with KV binding functions as learned linear attention rather than memorization.
Federated learning aggregation method (FedVG) using gradient guidance to address client drift and data heterogeneity.
Safety filtering framework for flow-based generative models with formal guarantees on constraint satisfaction.
Strategic risk aversion approach for training collaborative AI agents that generalize better with new partners.
Adversarial reinforcement learning dataset (AOT-SFT) to improve robustness of multimodal LLMs on visually complex scenes.
Maps failure regions in LLMs using MAP-Elites quality diversity search to characterize unsafe behaviors and vulnerabilities.
Energy-based theory for detecting concept drift in ECG signals, distinguishing physiologically plausible variation from true distribution shift.
Regularized online RLHF for Nash Equilibrium identification with generalized bilinear preferences modeling intransitive preferences.
ParamMem augments language agents with parametric reflective memory to improve reasoning through diverse self-reflection.
FinBloom presents a knowledge-grounding approach for LLMs to handle real-time financial queries with live data integration.
Apple proposes general active perception via reinforcement learning to handle uncertainty in partially observable robotic environments.
REA-RL uses online reinforcement learning with reflection to reduce overthinking and inference costs in large reasoning models.
CoMind introduces MLE-Live, a framework for evaluating LLM agents that engage with research communities in ML engineering tasks.
Uses Explainable Boosting Machines to identify overshooting tops in satellite imagery for weather forecasting, emphasizing interpretability in high-stakes ML.
pFedMMA proposes personalized federated fine-tuning with multi-modal adapters for vision-language models like CLIP on decentralized heterogeneous data.
20 fine-tuned Llama-3.1 8B variants specialized for High-Energy Physics, trained on arXiv abstracts with comparative domain analysis.
Research method for LLM alignment using meta-weighted online sampling to address distribution mismatch between offline data and evolving model policy.
DataMind framework trains scalable generalist data-analytic agents on diverse data formats for multi-step reasoning in scientific discovery.
Study analyzing how language models learn context-free grammar structures and subgrammars during pretraining.
FLOP algorithm performs causal structure discovery for linear models using fast parent selection and discrete search optimization.
Supervised Reinforcement Learning framework combines expert trajectories with step-wise reasoning to improve multi-step reasoning in open-source LLMs.
Contrastive weight steering method edits LLM parameters via weight arithmetic for post-training behavior modification without expensive retraining.
Equivariant neural networks improve RANS turbulence model closures using structure tensors for better generalization in fluid dynamics.
NuBench open benchmark for deep learning-based event reconstruction in neutrino telescopes addressing inverse problems in particle physics.
Non-IID sampling framework for flow matching models to reduce variance in expectation estimation with limited sampling budgets.
MEDIC applies machine learning for automated data quality monitoring and anomaly detection in particle physics collider experiments.
AutoSpec framework automatically refines logical specifications for reinforcement learning agents to improve policy learning from under-specified tasks.
QKAN-LSTM combines quantum-inspired methods with Kolmogorov-Arnold networks for improved sequential modeling with reduced parameter redundancy.
Tutorial on using differentiable programming frameworks like PyTorch and JAX to learn and design optimization algorithms automatically.
GreenServ proposes dynamic routing framework for efficient multi-model LLM inference using context-aware model selection to reduce energy consumption.
Research on training humanoid robots with reinforcement learning to handle diverse embodiments and complex behaviors through distillation methods.
SpatiaLab benchmark evaluates vision-language models' spatial reasoning capabilities on real-world tasks with visual noise and diverse spatial relationships.
SpatiaLab: benchmark evaluating vision-language model spatial reasoning on complex real-world visual scenarios.
Flow-based intermediate representation for few-shot robot imitation learning from human video demonstrations.
Property-preserving kernel operator learning for incompressible Navier-Stokes flow simulation and surrogate modeling.
SceneTok: tokenizer encoding 3D scene views into compressed, permutation-invariant tokens for diffusion models.
Theoretical analysis of stochastic mirror descent convergence with matrix parameters in overparameterized regime.
Janus-Q uses event-driven hierarchical reward modeling with textual signals for financial market trading.
Aletheia, Gemini 3 Deep Think-powered math research agent, autonomously solved 6 of 10 FirstProof challenge problems.
Research on context design for LLM-based probabilistic forecasting in mention prediction markets.
Framework for multi-level causal embeddings enabling mapping of detailed models into coarser causal model sub-systems.
veScale-FSDP improves fully sharded data parallel training flexibility for structure-aware methods and advanced optimizers.
Enveil is a self-hosted encrypted environment variable manager replacing .env files with runtime injection for secure secret management.