Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
Systematic investigation of vision encoder-LLM alignment in Vision-Language Models using Gromov-Wasserstein distance for principled model selection.
Systematic investigation of vision encoder-LLM alignment in Vision-Language Models using Gromov-Wasserstein distance for principled model selection.
Segment-Aligned Policy Optimization (SAPO) for LLM reinforcement learning that aligns credit assignment with reasoning step structure in multi-modal tasks.
Vision-language models with active reasoning via sequential Bayesian decision-making. VLM improvement through adaptive visual perception.
Multi-agent debate for on-policy distillation in agentic tasks. Teacher-student learning framework with agent trajectory optimization.
Framework for analyzing concept representations in neural networks via linear subspaces. Interpretability research applicable to model understanding.
Using RL to improve MLLMs on imbalanced regression tasks via distributional awareness. LLM training methodology addressing long-tailed distributions.
Framework for composing and auditing LoRA adapters from open pools for tasks. Parameter-efficient fine-tuning and model composition.
LLM coding agents with persistent memory, RAG, and RL feedback for software engineering. Architecture for agent memory and tool use.
Brain-inspired spiking neural network with time-delayed coordination for learning. Neuroscience-focused theoretical work on oscillatory dynamics.
Causal discovery method with per-edge trust scores for heterogeneously reliable external priors from diverse sources.
Bayesian generative modeling framework for missing data imputation with uncertainty quantification.
Federated learning approach for RF jamming detection in 5G networks preserving privacy vs. centralized learning.
Lossless KV cache compression technique for efficient disaggregated LLM serving with reduced transfer bottleneck.
Analysis of compliance gap in AI agents: verbal agreement vs. actual behavior divergence in following explicit instructions.
Vision-Language-Action model with anticipation-based subgoal generation for long-horizon embodied robotic tasks.
RL-based UAV navigation combining control Lyapunov functions and barrier functions for safe autonomous flight.
Multi-agent reinforcement learning framework using multi-step causal influence extraction for improved agent coordination.
Communication-efficient multi-agent control framework for remote action generation with bandwidth constraints.
Multimodal LLM approach for understanding dense charts with fine-grained visual reasoning and focus-driven mechanisms.
Extreme value theory applied to machine learning for extrapolation, regression, and anomaly detection in data-sparse regimes.
Hardware-software co-design for Vision Mamba inference on FPGA with quantization optimization.
Meta-learning framework for LLMs using hypernetworks and adaptive gating. LLM architecture innovation.
LLM-based entropy coding for text transmission over fixed-rate channels. Combines compression with neural networks.
Framework for standardizing randomized controlled trials in AI evaluation. Establishes RCT best practices.
Benchmark platform for multi-agent reinforcement learning with mixed cooperation/competition. Open research environment.
LLM-based framework for essay scoring using pairwise comparison transfer learning. LLaMA fine-tuning application.
Vision-language alignment method for ultrasound images using contrastive learning. Medical imaging foundation model.
Parameter-free optimization algorithm achieving O(ε^-5/3) complexity for non-convex functions. Theoretical ML research.
Multi-agent LLM framework with planner, actor, and memory manager roles for long-horizon planning and complex task automation.
Multi-agent framework with planner, actor, and memory manager roles for long-horizon LLM-based task automation.
Privacy-preserving multi-camera object detection framework using diffusion-based domain adaptation.
Information-theoretic analysis of Pearl's causal hierarchy quantifying description complexity across causal inference levels.
AI-generated examples of graph dominating sets using transformer-based reinforcement learning tool PatternBoost.
Systematic study of metric reliability in Vision-Language Model unlearning for GDPR compliance.
Multimodal ML framework for pneumonia screening combining symptoms, respiratory patterns, and imaging.
Study of perturbation effects in recursive LLM loops examining context-update rules (append, replace, dialog) and persistence of redirected behavior.
Anon optimizer improves adaptive methods like Adam by adjusting pre-conditioner adaptivity to generalize better across diverse optimization landscapes.
Explores application of causal discovery algorithms to automated legal argument generation, applying Pearl's causal reasoning methods.
ANO framework unifies PPO and SPO by establishing trust region approach that balances gradient information retention with optimization stability in deep RL.
FedPLT addresses federated learning challenges by partial layer training to reduce communication/computation overhead and handle device heterogeneity.
Decoding-time debiasing via process reward models reduces social biases in LLM outputs without model weight access or fine-tuning.
Study of structured output reliability gap in 7-9B language models, evaluating joint accuracy of mathematical correctness and JSON format compliance across prompting strategies.
SCHEMA evaluation framework assesses metacognitive stability of frontier AI models under adversarial pressure, studying cognitive collapse as safety failure mode.
FitText framework enables dynamic tool retrieval in AI agent reasoning loops by embedding retrieval directly in execution, addressing semantic gaps between task descriptions and API documentation.
Power sampling decoding method to efficiently locate multi-step solutions in LLM reasoning by biasing toward high-probability modes without additional training.
Goal-conditioned reinforcement learning approach combining graph neural networks for middle-mile logistics routing.
Deep reinforcement learning study using procedural map generators to improve navigation policy generalization across diverse environments.
Empirical study examining how task horizon length affects training dynamics for LLM-based agents solving extended interaction sequences.
Agentic research system reproduces NLP study on LLM style editing in 3 hours with human-in-loop, demonstrating rapid empirical research workflow.
Framework using AI agents to automate reproducibility assessment in scientific peer review, formalizing it as structured reasoning over research artifacts.