SHIFT: Robust Double Machine Learning for Average Dose-Response Functions under Heavy-Tailed Contamination
SHIFT: Robust double machine learning estimator for dose-response functions handling heavy-tailed contamination with Welsch loss.
SHIFT: Robust double machine learning estimator for dose-response functions handling heavy-tailed contamination with Welsch loss.
RSAT: Method training small LMs (1-8B) to produce step-by-step table reasoning with cell-level citations using structured JSON and GRPO.
Unified perspective on fine-tuning and sampling diffusion and flow models via stochastic optimal control and thermodynamics.
Information-geometric framework for adaptive sampling in graph diffusion models using Riemannian manifold geometry.
Analysis of Mamba's state space representations for extracting sentence embeddings without fine-tuning or pooling heads.
Research on how LLMs detect out-of-distribution inputs, revealing confounds in OOD detection methods like CED and RAUQ.
ML models for intrusion detection in intelligent transport systems using edge computing and AI for security.
Uses LLMs with reasoning to process sandbox behavior reports for improved PE malware detection beyond static features.
Continuous benchmark measuring AI inference at endpoint granularity across energy, speed, and latency dimensions.
Method for self-calibrating vision-language models online to reduce hallucinations without relying on external model supervision.
Production infrastructure system enabling feature efficiency rollouts without model retraining in large-scale ranking systems.
Analyzes attractor basin geometry and storage capacity limits in kernel Hopfield networks for associative memory.
Token pruning method for DeepSeek-OCR vision-language model to reduce inference costs while preserving text fidelity.
Proposes structured recurrent spiking neural networks with sparse connectivity and scalable learning without backpropagation.
On-chain benchmark for evaluating AI forecasting agents with permissionless evaluation resistant to overfitting and data contamination.
Fine-tunes small 3-4B parameter LLMs with LoRA for radiology tasks, enabling CPU deployment in resource-constrained environments.
CleanBase system detects malicious documents in RAG knowledge databases to prevent prompt injection attacks that compromise LLM integrity.
End-to-end autoregressive image generation pipeline jointly optimizing visual tokenizer and generative model using foundation models.
SAGA proposes workflow-atomic scheduling for AI agent inference on GPU clusters, treating entire agent workflows as first-class objects to reduce latency.
Tempus framework for efficient GEMM acceleration on AMD Versal edge hardware targeting LLM inference deployment constraints.
Four jailbreak attacks exploiting vision modality of vision-language models to bypass safety alignment through visual encoding and substitution.
BlenderRAG system uses retrieval-augmented generation to improve LLM code synthesis for 3D object generation in Blender with reduced errors.
EASE framework for federated multimodal unlearning that disentangles forgotten knowledge across modalities and decentralized clients.
Position paper arguing that agentic AI system orchestration layers should use Bayesian approaches for decision-making under uncertainty with tools and experts.
Themis introduces reward models for multilingual code generation that extend beyond execution feedback to enable multi-criteria scoring and policy alignment during LLM post-training.
LightKV method reducing KV cache memory overhead in large vision-language models by exploiting token redundancy.
Framework learning PDE solution operators by composing modular components conditioned on system regime features.
Federated learning approach for training diffusion models collaboratively while addressing data privacy and computational constraints.
Deep learning methods combining neural networks with numerical differential equation solvers for dynamics discovery and parameter estimation.
Federated learning approach for handling noisy labels across heterogeneous clients with different label noise types and data distributions.
Empirical study comparing exploration-exploitation tradeoff behavior of LLMs versus humans in multi-armed bandit experiments.
Generative modeling approach for random fields from limited training data using deep generative models.
Privacy amplification analysis for zeroth-order optimization in differentially private LLM fine-tuning with memory constraints.
Disentangled Safety Adapters framework for efficient AI safety through lightweight adapters decoupled from task models.
Diffusion model framework for solving inverse problems using piecewise guidance from conditional distributions.
Study of spontaneous deception in LLMs on benign prompts without explicit hidden objectives, examining trustworthiness risks.
PyFair framework for verifying individual fairness in DNNs using concolic testing with formal guarantees.
Optimal decision tree algorithms addressing expressivity and scalability through non-axis-parallel splits.
Diffusion-based framework for learning physical dynamics from incomplete and irregularly sampled observational data.
Selective prediction layer using MC-Dropout for knowledge tracing models to defer uncertain predictions to humans.
Method for tracking evolutionary relationships between LLMs via functional DNA-inspired representations without tokenizer assumptions.
DIANO framework for interpretable latent space modeling in scientific machine learning using neural operators.
Adaptive feature selection approach for graph neural networks that identifies and removes unnecessary features during training to improve interpretability.
Score-based greedy search method for identifying causal structure in partially observed linear systems, addressing challenges in causal discovery.
Byzantine-robust decentralized federated learning using sketch-based screening to reduce communication costs in privacy-preserving collaborative training.
Unified framework for iterative neural network compression combining pruning, quantization, and low-rank decomposition while mitigating accuracy degradation.
MemoryBench benchmark for evaluating memory and continual learning in LLM systems as alternative to scaling parameters and data.
Last-iterate convergence analysis of FTRL algorithm with Tsallis entropy in stochastic multi-armed bandits.
SynQuE formalizes synthetic dataset quality estimation using proxy metrics to rank datasets without annotations for data-scarce settings.
Uncertainty modeling approach for Real-Time Auction interception with distillation acceleration for traffic quality estimation.