KLCF framework addresses hallucination in LLM long-form generation by incorporating knowledge-level consistency into RLHF, aligning model outputs with its own knowledge boundaries.
Unsupervised framework using optimal transport for learning procedures from instructional videos by handling noise and execution variability.
Framework for automatically searching the internet to construct challenging benchmarks at scale without human curation for model evaluation.
Proof-of-concept combining role-playing games with LLM analysis to elicit moral profiles of users in software requirements engineering.
Theoretical analysis of Reinforcement Learning with Verifiable Rewards (RLVR) for post-training LLMs using binary feedback, introducing Gradient Gap metric.
Studies inference-time scaling using pause tokens to improve foundation model expressivity while maintaining parallelizability for faster reasoning.
Comprehensive review of Kolmogorov-Arnold Networks as structured alternatives to MLPs, covering theory, relationships to classical methods, and applications.
Proposes AsyncVLA, a Vision-Language-Action model using asynchronous flow matching for improved long-horizon robotic task execution with self-correction capabilities.
Stellar VLA: continual learning framework for vision-language-action models that evolves skill knowledge without parameter expansion.
RobustSora: benchmark for detecting AI-generated videos, controls for watermarks to isolate genuine generation artifacts.
SoccerMaster: vision foundation model for soccer understanding tasks from detection to semantic reasoning.
MediEval: benchmark linking EHRs to medical knowledge base for evaluating LLM reasoning and reliability in medical applications.
Deep Bayesian RL framework with learnable basis functions for improved generalization in Meta-RL tasks.
Research mapping human anti-collusion mechanisms to multi-agent AI systems. Studies how autonomous agents develop collusive strategies.
IGBO framework trains interpretable ML models balancing accuracy and explainability using feature importance hierarchies and gradients.
CSMCIR method for composed image retrieval combining text and images using CoT-enhanced alignment and memory bank.
HERMES: training-free architecture for efficient streaming video understanding in MLLMs using hierarchical KV cache memory.
Study showing LLMs are state-blind: ignore contextual/situational factors while capturing trait-based personas. Introduces Chameleon dataset.
Research on ensemble methods for causal discovery algorithms with expert guidance. Machine learning methodology for practical applications.
Cap-and-trade regulatory framework proposal to incentivize AI efficiency and accessibility, addressing resource equity and sustainability.
LLM-AutoDP framework using LLM agents to automatically process and clean domain-specific data for model fine-tuning.
Gossip-based decentralized learning algorithms for edge devices with communication efficiency, robustness to corruption, and low memory.
Theoretical analysis of optimal Attention/FFN ratios in disaggregated LLM serving architectures for efficient resource provisioning.
Meta-evaluation framework enabling self-evolving LLMs for non-verifiable tasks through LLM-as-Judge with quality-aware training.
FIT to Forget method for robust continual unlearning in LLMs handling sequential privacy and copyright deletion requests.
Leviathan transformer architecture decoupling input embeddings and output projections via learned embedding vectorization.
Neural solver for vehicle routing problems using lifelong learning with continually drifting task patterns and limited training.
DialectLLM framework for generating multi-dialectal conversational data beyond Standard American English, addressing LLM dialect representation.
Analysis of forgetting illusion in concept erasure for diffusion models via latent variable optimization attacks.
Statistical membership inference method for reliably auditing machine unlearning in models, addressing right to be forgotten.
Workflow for systematic evaluation of audio description quality using human raters and vision-language models at scale.
Novel training method for time-series forecasting models incorporating autoregressive rollout and error-growth heuristics from LLM training.
Theoretical analysis of transformer capabilities on the PARITY task, examining fundamental computational limits of neural architectures.
Probabilistic framework formalizing code selection and generation paradigms for test-driven development with AI assistants, analyzing environment-interaction strategies.
Flow matching approach for diffusion-based robotic policies reducing inference latency through informed noise sampling.
Adaptive curriculum enhanced group relative policy optimization for autonomous ML engineering agents using RL.
Attribution evaluation framework for multimodal LLMs assessing grounding across heterogeneous sources and modalities.
Parallel reasoning approach for visual comprehension in LLMs shifting from depth to parallelism to improve exploration.
Framework integrating deep generative models with quantum annealing for molecular design beyond training data.
Transformer architecture rethinking temporal and channel dependencies in medical time series like EEG and ECG.
Confidence-gated reflection method for reward modeling combining interpretability and efficiency in LLM alignment.
Cross-modal study comparing human preferences in text vs. speech for preference-based reinforcement learning evaluation.
Position paper arguing LLM-based agents alone are insufficient for social simulation, identifying systematic limitations.
LLM-based embodied agent with personality traits for persistent autonomous behavior in dynamic environments.
Dynamic Chunking Diffusion Transformer replacing static patchification with adaptive token compression for visual generation.
Benchmark quantifying hallucinations in LLMs on medical textbook QA tasks with fixed evidence sources.
Meta-learning approach for knowledge editing improving accuracy-editability trade-off through connected two-stage optimization.
LLM-based approach generating span-level evidence annotations for ICD medical coding from clinical documents.
Contextual rubric reward framework for RL extending beyond scalar RLHF with multi-dimensional structured evaluations.
Interventional boundary discovery method enabling RL agents to identify controllable state dimensions from noisy observations.