Explicit Dropout: Deterministic Regularization for Transformer Architectures
Deterministic dropout formulation expressing regularization as additive loss terms for Transformer architectures with explicit optimization objectives.
Deterministic dropout formulation expressing regularization as additive loss terms for Transformer architectures with explicit optimization objectives.
CHASM dataset and benchmark for evaluating multimodal LLMs on detecting covert advertisements in Chinese social media moderation tasks.
Federated learning framework with differential privacy and privacy-preserving initialization for heterogeneous cross-device deployments.
Study of relationship between model calibration and loss surface curvature in neural networks, showing calibration emerges during training on vision tasks.
Occupancy reward shaping method extracting temporal information from generative world models to improve credit assignment in offline goal-conditioned reinforcement learning.
GRPO-VPS extends Group Relative Policy Optimization with verifiable process supervision to improve credit assignment for LLM reasoning tasks.
Systematic empirical study of transformer compression across 40+ experiments on GPT-2 and Mistral 7B identifying five structural properties relevant to model compression.
MGDA-Decoupled proposes geometry-aware multi-objective optimization for DPO-based LLM alignment balancing helpfulness, truthfulness, and harmlessness objectives.
COMPASS framework using parameter-efficient fine-tuning (PEFT) and adaptive semantic sampling to improve multilingual LLM performance and reduce cross-lingual interference.
Tokenised flow matching technique for hierarchical simulation-based inference to reduce simulator evaluation costs through likelihood factorization.
Proposes Supplement Generation Training (SGT), a method to train smaller LLMs to generate supplemental text for improving agentic task performance without expensive post-training.
Research on reinforcement learning with verifiable rewards (RLVR) exploring mixed-policy methods to improve convergence by combining off-policy and on-policy trajectories.
Simulation-based Bayesian inference framework for real-time industrial equipment condition monitoring that overcomes MCMC computational bottlenecks.
F²LP-AP combines label propagation with adaptive kernels for fast semi-supervised node classification on heterophilous graphs without expensive iterative training.
Federated continual learning approach for autonomous vehicle fleets that adapts to evolving terrain while preventing catastrophic forgetting across network layers.
ParetoSlider framework extends diffusion model post-training with multi-objective reward control, enabling inference-time trade-off adjustment between conflicting goals without early scalarization.
Stream-CQSA algorithm enables efficient long-context LLM attention computation by removing full tensor fitting requirement, reducing memory complexity from quadratic to near-linear.
FedSIR multi-stage federated learning framework using spectral client identification and relabeling for handling noisy labels in distributed training.
ZeroFolio feature-free algorithm selection method using pretrained text embeddings and k-nearest neighbors without domain knowledge.
Study on data augmentation and transformer-based models addressing class imbalance in automated scoring of student scientific explanations.
LLM approach for anti-money laundering transaction triage with explainable evidence retrieval and counterfactual verification for regulated workflows.
ThermoQA benchmark of 293 thermodynamics problems across three tiers evaluating reasoning capabilities in six frontier LLMs.
EvoForest machine learning paradigm using open-ended evolution of computational graphs for discovering transformations in structured prediction.
TTKV temporal-tiered KV cache system reducing memory footprint for long-context LLM inference using hierarchical precision approach.
Domain-specific LLM for tuberculosis care in South Africa developed to assist patients and healthcare providers.
Systematic study identifying precision-induced output disagreements in LLMs across floating-point and quantized formats affecting reliability.
SkillGraph system for LLM agent tool sequence recommendation using execution-transition graphs mined from 49,831 successful workflows.
Neural network approach for predicting molecular energies and forces incorporating temporal information from molecular dynamics simulations.
MIRROR benchmark with 250,000 evaluation instances across 16 LLMs measuring metacognitive calibration and self-knowledge in decision-making.
Large-scale empirical study of 830+ files across 12 LLM models examining how test code structure affects AI code generation quality.
Multi-agent LLM framework for RTL code generation achieving 95.9% functional correctness with validation-first approach and synthesis awareness.
Systematic mechanistic analysis of LLM quantization failure modes, distinguishing between signal degradation and computation collapse in 2-bit vs 4-bit quantization.
DistortBench: Diagnostic benchmark with 13,500 questions evaluating vision-language models on image distortion perception across 27 types.
Extends blicket detector paradigm to test AI agents' capacity for causal reasoning and hypothesis space restructuring through experimentation.
Self-play paradigm for LLM training using rubric-based reward models on pre-training text with minimal supervision for open-ended tasks.
SkillLearnBench: First benchmark for evaluating continual skill learning in LLM agents across 20 real-world tasks in 15 domains.
HiPO extends Direct Preference Optimization for LLMs with hierarchical feedback on reasoning steps, addressing limitations in complex task alignment.
Feature-scoped ML approach predicts query slot-time in cloud data warehouses before execution for cost estimation.
Framework for robust stochastic optimization under distribution shift using relevant source distributions for decision-making.
Meta-Tool empirically compares hypernetwork LoRA adaptation vs few-shot prompting for tool-use in small language models.
World model-based reinforcement learning approach for safe autonomous robotic endovascular intervention with variable anatomies.
WildFireVQA benchmark combines RGB and thermal imagery for visual question answering on aerial wildfire monitoring.
Mol-Debate uses multi-agent debate to improve structural reasoning for text-guided molecular design in drug discovery.
Reinforcement learning-based sample selection improves transfer learning in low-resource, imbalanced clinical settings.
AROMA system uses augmented reasoning over multimodal architecture for virtual cell genetic perturbation prediction.
Surrogate modeling framework to interpret black-box LLM knowledge and explain medical predictions quantitatively.
Systematic study of hallucination in AI models of fluid dynamics, showing physically implausible but visually realistic predictions.
WebGen-R1 uses reinforcement learning to train small LLMs to generate functional, multi-page websites end-to-end.
DialToM benchmark evaluates theory of mind reasoning in LLMs through human-verified dialogue trajectory forecasting tasks.
VTouch++ dataset for bimanual robotic manipulation with vision-based tactile sensing for contact-rich tasks.