A feature-stable and explainable machine learning framework for trustworthy decision-making under incomplete clinical data
Machine learning framework for trustworthy clinical decision-making with feature stability under incomplete data.
Machine learning framework for trustworthy clinical decision-making with feature stability under incomplete data.
Research on improving LLM-based recommendation systems using self-hard negatives from intermediate layers for better preference learning during fine-tuning.
Taxonomy and comparative study of uncertainty quantification methods for detecting hallucinations in long-form LLM outputs.
Study on generative-retrieval architectures in web search and how LLMs have transformed information retrieval practices.
arXiv paper on Jolt Atlas, a zero-knowledge ML framework for verifiable ONNX tensor operation inference using lookup arguments.
Audit of personal data associations in 8 LLMs using LMP2 privacy probe, examining how models retain and surface personal information.
arXiv paper on training neural networks with Boolean threshold functions where all node values are strictly ±1.
arXiv paper introducing LORA-CRAFT, a parameter-efficient fine-tuning method using Tucker decomposition on transformer attention weights.
arXiv paper identifying transformer attention heads functioning as membership filters, analyzing their spectrum of testing strategies across language models.
arXiv paper systematically evaluating mechanistic interpretability in single-cell foundation models using 37 analyses and 153 tests.
Position paper proposing AI co-design for autonomous particle accelerator operation with minimal human intervention.
arXiv paper introducing MASPO, a reinforcement learning method improving gradient utilization and probability mass handling for LLM reasoning.
arXiv paper analyzing how normalization strategies impact Transformer expressivity for time series representation learning.
arXiv paper on Deep-Flow, an unsupervised anomaly detection framework for autonomous vehicles using optimal transport conditional flow matching.
arXiv paper testing whether speech LLMs behave identically to ASR-to-LLM cascades across four models and six tasks.
arXiv paper on relevance-guided online meta-learning for geospatial discovery under resource constraints and dynamic environments.
arXiv paper on anytime-valid statistical watermarking for distinguishing machine-generated content from human text in LLMs.
arXiv paper on variance control in asynchronous off-policy RL for LLMs, addressing high variance from stale rollouts in critic-free methods.
arXiv paper analyzing weak vs strong verification mechanisms in LLM reasoning systems, examining cost-reliability tradeoffs in verification loops.
arXiv paper on Reverso, a time series foundation model for zero-shot forecasting that scales to hundreds of millions of parameters.
FAMOSE: ReAct-based agent for automated feature engineering in tabular data that autonomously explores and generates optimal features without domain expertise.
Black-box adversarial attack method on Large Vision-Language Models using fine-grained detail targeting to address gradient-free optimization challenges.
MARS framework for reward modeling using margin-aware training and self-refinement to reduce reliance on costly human-labeled preference data.
Novel pruning technique for Diffusion Language Models that optimizes inference efficiency by reconsidering attention sink preservation assumptions.
Research on embodied AI agents using LLMs for open-ended dialog to infer and accomplish diverse user goals efficiently and robustly.
GAI: multi-agent LLM framework with reflection and dialogue for collective reasoning to drive innovation.
Framework studying AI-assisted human decision-making where humans learn through repeated interactions with algorithms.
Method for learning user-specific reward models in RLHF to capture individual preferences in LLM training.
Evaluation framework for assessing health-focused LLMs on personalized response quality with scalable methodology.
Theoretical correspondence between bounded GNNs and first-order logic fragments characterizing expressive power.
∞-THOR: framework for long-horizon embodied AI tasks with benchmark testing long-context reasoning across extended trajectories.
SPECS: method for faster test-time scaling in LLMs using speculative drafts to reduce latency while maintaining performance.
Embodied AI system enabling autonomous drones to make adaptive decisions for sudden events using visual language models.
PROBE: benchmark for measuring proactive problem-solving in LLM agents across extended contexts and time horizons.
SCL: modular agent architecture separating cognition into five phases with soft symbolic control governance layer for LLM agents.
CaveAgent: framework converting LLM-as-text-generator to LLM-as-runtime-operator with dual-stream architecture for long-horizon task execution.
Offline multi-agent RL using local-to-global world models to enable conservative policies to generalize beyond dataset support.
AUTOBUS: neuro-symbolic AI system combining LLMs with deterministic logic for autonomous business process reconfiguration and execution.
SpikeScore: method for cross-domain hallucination detection in LLMs that generalizes across different domains better than existing approaches.
ADP-MA: framework for autonomous data processing using meta-agents that monitor, manage, and optimize end-to-end pipelines after deployment.
Federated learning framework addressing device heterogeneity and non-IID data with differential privacy using bi-level optimization.
EduEVAL-DB dataset for training AI tutors to evaluate pedagogical quality of educational explanations across K-12 subjects.
CoreCraft: enterprise RL environment for training generalizable AI agents in customer support simulation with 2,500+ entities and 23 tools.
Benchmark evaluating physical safety risks of LLMs controlling robotic systems, categorizing drone-related threats and harms.
Knowledge distillation pipeline to compress Dust3r foundation model for faster 3D reconstruction and visual localization.
Meta-RL approach using skill decomposition with improved robustness to noisy offline demonstrations for long-horizon tasks.
Rex reversible exponential Runge-Kutta solvers for neural differential equations in generative models.
Self-organizing maps combined with vision transformers to improve ViT performance on small datasets.
Certified backdoor defense for DNNs using sample-specific smoothing noise against training data poisoning attacks.
ReplaceMe training-free depth pruning method replacing transformer blocks with linear operations for model compression.