Experimental evaluation of Free-Market Algorithm orchestrated Mixture-of-Experts with cost-penalized fitness for domain adaptation.
Optimal decomposition technique for low-rank approximation of LLM weights enabling efficient fine-tuning and inference.
Method for language agents to optimize test-time adaptation policies through iterative refinement during inference.
Reinforcement learning approach with verification for iteratively improving LLM policies based on actual performance gains.
Framework for human-AI cooperation that models fatigue-induced performance degradation in learning-to-defer systems.
Method for verifiable repair of transformer vulnerabilities to adversarial perturbations with inner-layer guarantees.
Graph partitioning technique using embeddings to enable scalable distributed training of graph neural networks.
Transfer learning methodologies for Bayesian network structure learning with scarce data.
Model-based learning approach for finite-window policies in partially observable Markov decision processes.
Method for efficiently evaluating LLM downstream performance during training without expensive full inference.
Theoretical analysis of dependency networks using information geometry perspective for modeling complex systems.
Analysis showing how irrelevant context degrades LLM reasoning performance despite test-time scaling capabilities.
Generative model approach using adversarial distribution alignment to bridge simulation-to-experiment gap in scientific domains.
ORCA framework calibrating LLM sampling through conformal prediction to improve test-time reasoning efficiency and generalization.
Multiscreen mechanism for language models enabling absolute query-key relevance assessment beyond relative attention redistribution.
CliffSearch agent framework for scientific algorithm discovery combining LLM-guided search with structured evolution of theory and code.
Mathematical framework analyzing what determines forecast skill in AI weather prediction, emphasizing training methodology over architecture.
PhoneticXEUS model for robust multilingual phone recognition trained on large-scale data with pretrained representations.
LLM-based recruitment tool identifying requisition-specific competencies through dynamic few-shot prompting and reflection.
Text-based harmonization approach using LLMs to unify multi-institutional EHR data without explicit schema standardization.
LLM-based approach to identify enterprise architecture debt indicators from unstructured documentation in organizations.
Framework combining vision language models with RL for dense reward generation in long-horizon robotic tasks to reduce manual reward engineering.
GenoBERT uses transformers for reference-free genotype imputation without ancestry bias.
HIVE framework for hierarchical pre-training of vision encoders integrated with large language models for vision-language alignment.
Transformer-based models for detecting software vulnerabilities in C/C++ using program slices.
MambaVoiceCloning uses state-space models and diffusion for efficient text-to-speech synthesis without attention layers.
Studies grokking in feature learning kernels via Recursive Feature Machine, showing data symmetry breaking is necessary for generalization.
Reverse-engineers gpt-oss-20b tool definitions from in-distribution calls and builds native harmony agent harness with open-source implementation.
Proposes decision-centric framework separating control decisions (answer, retrieve, tool use) from LLM generation in agent architectures.
EgoNav: Humanoid robot navigation system trained on 5 hours of human walking data using diffusion models and frozen DINOv3 backbone.
Shapley-guided approach using derivative-free optimization to repair DNNs affected by backdoors, adversarial attacks, and unfairness.
Studies policy gradient methods for multi-agent reinforcement learning in partially observable Markov potential games.
CheXOne: Vision-language foundation model for chest X-ray interpretation with explicit reasoning about visual evidence.
Introduces Uni-SafeBench, a safety benchmark for unified multimodal large models testing both understanding and generation capabilities.
Studies trade-off between pretraining corpus size and retrieval-augmented generation for language models under fixed data budgets.
CircuitProbe predicts reasoning circuits in Transformers from activation statistics in under 5 minutes, achieving 3-4 orders of magnitude speedup over brute-force methods.
Benchmarks State-Space Models (Mamba) against Transformers and BiLSTM for historical newspaper OCR, addressing quadratic complexity limitations.
Stochastic Attention inspired by connectome topology provides linear-time expressive attention mechanism.
Study shows multimodal LLMs fail at detecting 3D spatial inconsistencies across multiple views.
PARE framework simulates realistic user interactions for evaluating proactive AI agents and assistants.
StanceMoE uses mixture-of-experts for actor-level stance detection in geopolitical texts.
Dataset and analysis of autonomous coding agent contributions to real-world GitHub projects over time.
MyPhoneBench evaluates privacy compliance of mobile phone-use agents completing benign tasks.
Query-conditioned evidential keyframe sampling for efficient multimodal LLM-based long-form video understanding.
MoA-DepthCLIP adapts CLIP vision-language model for monocular depth estimation with parameter-efficient adapters.
PaperRecon framework evaluates quality and hallucination risks in papers generated by AI coding agents.
NARCBench for detecting multi-agent collusion using multi-agent interpretability on LLM agent activations.
S0 tuning zero-overhead adaptation of hybrid recurrent-attention models outperforming LoRA on code generation.
RELISH lightweight architecture for text regression with LLMs using iterative latent state refinement.
Survey on Graph Neural Network acceleration techniques across algorithms, systems, and customized hardware.