Agentic Exploration of Physics Models
Agentic framework for automated scientific discovery that iteratively explores unknown systems through experiments and analysis without domain-specific tailoring.
Agentic framework for automated scientific discovery that iteratively explores unknown systems through experiments and analysis without domain-specific tailoring.
Framework for learning abstract world models that jointly represent symbolic states and causal processes for endogenous and exogenous dynamics in robot planning.
Framework converting internet videos of human computer use into training data for computer-using AI agents via UI trajectory extraction.
Foundation model approach to auto-bidding in online advertising that generalizes across different bidding scenarios.
EcoAlign framework balances safety, utility, and computational cost in aligning Large Vision-Language Models against jailbreak attacks.
Echo-CoPilot: agentic framework combining multi-perspective workflow with knowledge-graph guidance for reliable echocardiography interpretation.
Reason2Decide: two-stage training framework for clinical decision support LLMs to generate predictions with self-aligned explanations.
MultiSessionCollab benchmark and method for long-term conversational agents to learn and leverage user preferences across multiple sessions.
Position paper arguing agentic evolution via deployment-time adaptation is needed to close the train-deploy gap in LLM systems.
Method for compiling random forest classifiers into circuits for explainability and tractable computation of complete generalizations.
Formalizes causal Rung Collapse where LLMs learn spurious associations instead of causal relationships, proposes epistemic regret minimization solution.
Study showing fine-tuning vision-language agents on narrow harmful datasets causes emergent misalignment generalizing across unrelated tasks and modalities.
BAPO: off-policy reinforcement learning framework improving data efficiency in LLM post-training by selecting diverse training experiences.
Aletheia mathematics research agent solved 6 of 10 FirstProof challenge problems autonomously using Gemini 3 Deep Think reasoning.
Framework for decision-level evaluation of AI agents in AutoML pipelines beyond outcome metrics, assessing intermediate reasoning steps.
Survey and framework for personalized LLM-powered agents that adapt to individual users over extended interactions with evaluation methods and research directions.
Human study measuring whether LLM access improves novice performance on biology tasks versus internet-only baselines, with dual-use risk implications.
EMPA framework evaluates how well LLM dialogue agents maintain persona-aligned empathy across multi-turn conversations using process-oriented metrics.
WebChain: 31,725 human-annotated web interaction trajectories with 318k steps in multi-modal format for training and evaluating web agents.
LLMs used to synthesize executable game design patterns from high-level gameplay ideas, focusing on goal patterns and structural constraints in game creation.
Framework for automated frontier AI risk evaluation using executable code environments and LLM-based simulators.
Benchmark evaluating LLM ability to generate interactive HTML-based MiniApps with dynamic interfaces and logic.
Study showing reasoning and deliberation increase honesty in LLM responses on moral trade-off scenarios.
Framework for efficient LLM distillation that focuses training on problems at frontier of student capability.
Protocol for detecting self-preservation behaviors in autonomous agents to distinguish intrinsic from instrumental objectives.
Framework coordinating multiple LLM-based agents through verification loop for complex query resolution with DAG decomposition.
Benchmark with 6,372 multimodal reasoning instances that evaluates LLM reasoning transparency through verifiable intermediate steps.
Research on semantic invariance property of LLM-based autonomous agents under input variations to ensure stable reasoning.
Method using counterfactual thinking to identify and address bias and fairness issues in machine learning models.
Research on automated prompt generation and optimization techniques for improving LLM performance through meta-prompting approaches.
Ayn: domain-specific tiny language model pretrained from scratch for Indian legal NLP tasks as alternative to large LLMs.
Survey examining computerized adaptive testing through machine learning lens, covering personalized assessment methods across domains.
Universal approximation theorem and operator learning methods for continuous nonlinear operators in Banach spaces using orthogonal projections.
TraffiDent dataset aligning traffic dynamics and incident data across 16,972 nodes for understanding their interplay.
Analysis of how skip connections in deep networks enhance adversarial example transferability across models.
Time series forecasting approach accounting for latent confounders using causal inference to improve prediction accuracy.
Causal inference method using LLMs to quantify effects of textual interventions on social systems from observational data.
VisionZip reduces computational costs in vision-language models by compressing redundant visual tokens while maintaining performance.
RRNCO addresses real-world deployment of neural combinatorial optimization for vehicle routing by handling asymmetric costs and edge-based features.
Review of LLM-driven approaches for creating virtual agents with personality in VR environments using multimodal outputs.
Mask Fine-Tuning (MFT) introduces a novel LLM fine-tuning method that improves performance by selectively masking model components without updating weights.
MegaScale-Data addresses computational challenges in training large foundation models from multiple data sources by optimizing dataloader distribution across parallel ranks.
Credit assignment method (QLLM) for multi-agent RL eliminating predefined mixing networks through improved value decomposition and interpretability.
Nemotron-CrossThink extends RL-based self-learning from math reasoning to broader domains using verifiable reward structures and diverse tasks.
PCCL library for performant collective communication in distributed AI training on GPU supercomputers, addressing NCCL limitations.
Aitomia platform combining LLM-based agents and chatbots to assist with atomistic and quantum chemical simulations setup and analysis.
VideoSafetyEval benchmark with 11.4k video-query pairs across 19 risk categories for evaluating and defending Video LLM safety.
Method for improving LLM reasoning without expensive RL or high-quality demonstrations using weak supervision and incentive signals.
Inference-time alignment method for LLMs that searches in continuous response space using reward models for improved exploration.
SVD-based compression method (ERC-SVD) for efficient LLM deployment with error control and low-rank approximation techniques.