MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
MixAtlas: data mixture optimization method for multimodal LLM training with uncertainty-aware reweighting.
MixAtlas: data mixture optimization method for multimodal LLM training with uncertainty-aware reweighting.
PolyBench: multimodal benchmark for evaluating LLM forecasting and trading on live prediction market data.
Formal methods approach to verified neural network explanations with mathematical guarantees. Combines XAI with formal verification for safety-critical domains.
CROP: Automatic prompt optimization method with regularization for token efficiency in LLM reasoning. Balances accuracy against inference cost and latency.
PriHA: RAG-enhanced LLM framework for primary healthcare assistant in Hong Kong, grounding clinical guidelines to improve access and reduce healthcare costs.
Neuro-Oracle: Trajectory-aware agentic RAG framework for epilepsy surgical prognosis using 3D Siamese contrastive encoding and longitudinal MRI analysis.
Knowledge graph RAG with agentic recursive crawling for enterprise documents. Addresses hierarchical information capture in complex document ecosystems.
Tier-based adaptive routing framework for hybrid RAG across financial, legal, medical documents. Compares vector-based agentic vs hierarchical reasoning approaches.
TRACE: Multi-agent LLM framework using counterfactual explanations for sustainable tourism recommendations. Modular orchestrator-worker architecture.
FRESCO benchmark for evaluating re-rankers in RAG systems under evolving semantic information. Addresses temporal staleness in LLM-grounded retrieval.
Analysis of Claude Code agentic architecture via TypeScript source code, compared with OpenClaw open-source AI agent system. Identifies design patterns and human values.
Shapley value-based ensemble learning for explainable fraud detection with regulatory compliance. ML research on interpretability in financial domain.
Optimistic policy learning under adversarial exogenous factors with regret and constraint violation guarantees.
Counterfactual routing method to mitigate hallucinations in sparse Mixture-of-Experts models by activating dormant experts on long-tail knowledge.
Framework for evaluating agent performance under simulated marketplace dynamics with user switching, routing, and competitive constraints.
ReviewGrounder: rubric-guided, tool-integrated LLM agents for generating substantive peer review feedback grounded in existing work.
GUI-Perturbed framework reveals brittleness in GUI grounding models through controlled perturbations of visual scenes and instructions.
Value gradient flow method for behavior-regularized reinforcement learning in LLM finetuning without reparameterized policy gradients.
Contribution weighted group relative policy optimization for training LLM-based search agents with improved credit assignment and value estimation.
Tensor networks from quantum physics integrated into machine learning models for compressed representations of complex correlations.
EuropeMedQA: multilingual, multimodal medical examination dataset for evaluating LLM performance on non-English diagnostic tasks.
DharmaOCR: specialized small language models for structured OCR with benchmark covering printed, handwritten, and legal documents.
Study of agentic LLM systems for binary reverse engineering, examining limitations in obfuscation, timing, and unique architecture handling.
Attribution guidance method to improve faithfulness of textual explanations from LLMs, reducing gap between convincing and accurate rationales.
Uses Mamba SSM and LLM chain-of-thought reasoning to filter confounders in biomarker discovery from RNA-seq data.
Sample complexity bounds for best-arm identification in autonomous reasoning with safety guarantees for LLM-based node expansion in planning.
APEX-MEM: conversational memory system for LLMs using property graphs with temporal reasoning and entity-centric framework to improve long-term context retention.
Analysis of modal dependence in multimodal language models showing text centroid structure is 4x more critical than visual structure for performance.
SatBLIP: satellite-specific vision-language model for rural context understanding and social vulnerability index prediction from satellite imagery.
Modular continual learning architecture addressing catastrophic forgetting through task-specific experts, gatekeeper routing, and simultaneous pipeline training.
Step-level diffusion model alignment approach using reinforcement learning to balance multiple human preference objectives beyond single reward optimization.
First formal framework analyzing coalition formation in multi-agent LLM systems using hedonic game theory with stability guarantees and convergence proofs.
BiCon-Gate system for dialogue fact-checking using de-colloquialization and consistency gates to handle informal language in multi-turn conversations.
SpaceMind: modular vision-language agent framework for autonomous orbital servicing using skill modules, dynamic routing, and Model Context Protocol tools.
Research examining how LLMs rely on shallow heuristics and memorization rather than genuine reasoning in automated test generation for software like SAP HANA and LevelDB.
HRM-LM: Empirical study of hierarchical shared-weight recurrence versus independent Transformer layers for language model representation.
Theoretical analysis of multi-layer SSM expressiveness, showing fundamental compositional limitations and CoT benefits.
PyTorch toolkit for news recommendation research supporting learners with conceptual and practical experience.
Agent communication language with constrained extensibility enabling safe heterogeneous agent coordination across domains.
Multi-agent hierarchical framework using LLMs to generate synthesizable Verilog for large hardware designs with structural reasoning.
Corpus2Skill distills document corpora into navigable hierarchical skill directories for LLM agents performing QA and RAG tasks.
Framework leveraging LLM outputs as auxiliary data for operations management and causal inference under distribution shift.
Shared log system enabling AI agents to reason over streaming data without performance interference in data-streaming architectures.
Reverse-engineering framework to mechanistically decode complex emotional and cognitive constructs within LLM internals.
Framework using causal intervention on attention heads to reduce toxic content generation in LLMs while maintaining quality.
Security vulnerabilities in large audio-language models exposed through imperceptible adversarial audio injection attacks.
Retrieval-augmented approach for automating clinical value set authoring by grounding LLM generation in standardized medical vocabularies.
Studies clarification question generation in software engineering tasks, identifying which missing information types most affect task success with LLM assistants.
ELMoE-3D optimizes Mixture-of-Experts model serving via speculative decoding and memory-centric architectures for on-premises deployment.
StoryCoder framework transforms fragmented problem conditions into coherent narratives to improve LLM code generation through better structured reasoning.