Insect-inspired modular architectures as inductive biases for reinforcement learning
Insect-inspired distributed neural architectures as inductive biases for continuous control reinforcement learning.
Insect-inspired distributed neural architectures as inductive biases for continuous control reinforcement learning.
Training method using weak supervision to prevent sandbagging where capable models underperform verification ability.
Generative AI methods for creating synthetic malware samples with obfuscation techniques for security research.
Reinforced Iterative Classification framework allowing variable compute budgets and refined predictions via iterative refinement.
PermaFrost-Attack: poisoning attack via stealth pretraining seeding to plant backdoors in LLM training data.
Statistical framework for multi-agent LLM systems to detect unreliable decisions in self-harm risk screening.
Sum-of-Checks framework decomposing vision-language model reasoning for reliable surgical safety assessment.
Framework for estimating tail risks and rare harmful outputs in LLM distributions at population scale deployment.
Dynamic routing technique for offline reinforcement learning that balances Q-value improvement with dataset support constraints.
Research on how LLMs detect and correct their own errors using second-order confidence signals from decision neuroscience.
FETS Benchmark: Foundation models outperform dataset-specific approaches for generalizable energy time series forecasting.
Japanese medical foundation model balancing scaling laws and task-specific efficiency for clinical risk prediction on longitudinal data.
SOC-ICNN: Input convex neural network architecture generalizing from linear programming to second-order cone programming.
Extension of neural activation coverage technique for uncertainty estimation in regression tasks of pre-trained neural networks.
Hidden failure modes of gradient modification under Adam optimizer in continual learning with adaptive decoupled moment routing solution.
Analysis of distance-misaligned training failure modes in graph transformers with adaptive control mechanism.
HubRouter: Sub-quadratic routing module replacing O(n²) attention with O(nM) hub-mediated routing for hybrid sequence models.
FeatEHR-LLM: Framework using LLMs for automated feature engineering in electronic health records with irregular temporal patterns.
SOLAR-RL: Semi-online reinforcement learning framework for training multimodal LLM-based GUI agents on complex navigation tasks.
Data-free contribution estimation in federated learning using gradient von Neumann entropy to identify client importance without privacy leakage.
SpikingBrain2.0: 5B brain-inspired foundation model using spiking neural networks for efficient long-context inference with reduced training overhead.
Proposes adaptive head budgeting mechanism for multi-head attention in Transformers to improve efficiency by selectively activating attention heads based on task requirements.
Self-supervised learning approach using action-conditioned world models for cardiac disease detection, shifting from invariance-based to dynamics-aware objectives.
Research evaluating Shapley value variants for explainable AI in high-stakes applications, comparing theoretical formulations and human utility alignment.
WG-SRC white-box probe diagnosing graph neural network mechanisms and dataset feature requirements for node classification.
Transformer model extracting historical lexical structure and cognates from modern Bantu language morphological data.
Active experiment selection method for budget-aware scaling law fitting, reducing costs of planning large-scale ML training runs.
Evaluation framework testing whether LLMs perform genuine mathematical reasoning vs pattern matching on novel abstract mathematical problems.
Code generation pipeline study showing execution feedback outweighs topology complexity in 1-3B model composition on HumanEval.
MambaCSP: Hybrid state-space model combining attention and Mamba for hardware-efficient channel prediction with subquadratic scaling.
RE-CONFIRM framework evaluating robustness of brain biomarkers extracted by foundation models from dynamic functional connectivity data.
LLM prompt sensitivity investigation comparing instruction-based vs example-based prompting through shared lexical task representations.
Lightweight RAG and LLM approach for clinical patient-trial matching over heterogeneous EHR data with improved scalability.
LoRA adapter placement study in hybrid language models combining attention and recurrent components (Qwen, Falcon architectures).
Sovereign Agentic Loops: Control-plane architecture decoupling LLM reasoning from execution to improve safety in API-calling agents by emitting structured intents instead of direct outputs.
ArXiv paper on unsupervised anomaly detection for retinal OCT imaging without labeled data. Medical imaging ML research.
ArXiv paper on multi-armed bandits optimizing statistical utility functionals via influence-function gradients. Theoretical ML research.
ArXiv paper on adaptive control for constrained linear quadratic regulator achieving optimal regret bounds. Control theory and learning research.
ArXiv paper using multimodal diffusion models to enhance polarized light and EBSD microscopy data. Domain-specific ML application.
ArXiv paper on generative AI for marine propeller design using physics-based data generation. Applied ML research for engineering design.
ArXiv paper analyzing benchmark hacking in ML contests—gaming evaluation metrics without genuine improvement. Research on evaluation methodologies.
ArXiv paper on learning-augmented robotic control for manufacturing with real-world production validation. Applied ML research for industrial automation.
ArXiv paper on algorithmic feature highlighting for human-AI decision-making systems. Research on interpretable AI agent design.
ArXiv paper on LUT-aware neural network training for FPGA inference with ultra-low latency. Hardware-efficient ML architecture research.
ArXiv paper on pliable rejection sampling with kernel estimators for sampling difficult distributions. Machine learning research with new approach.
ArXiv paper on adaptive dictionary learning for kernel ridge regression to reduce space complexity. Machine learning research addressing scalability.
ArXiv paper on Conformalized Super Learner ensemble method with interval prediction uncertainty quantification. Machine learning research with theoretical contributions.
Research formalizing hidden randomness in LLMs via background temperature concept, analyzing nondeterminism from batch variation, kernel issues, and floating-point arithmetic.
Empirical evaluation of collective intelligence in large-scale LLM agent societies using MoltBook platform with 2M+ agents, testing emergence of intelligence at scale.
SnapLog approach extracts event data from video streams using image embeddings for process mining and business process management applications.