Critical Challenges and Guidelines in Evaluating Synthetic Tabular Data: A Systematic Review
Systematic review of 134 studies evaluating quality and reliability of synthetic tabular health data generation across 2067 relevant papers.
Systematic review of 134 studies evaluating quality and reliability of synthetic tabular health data generation across 2067 relevant papers.
Research on silent neuron theory and plasticity preservation in deep RL for adaptive video streaming under heterogeneous network conditions.
Autofocus Retrieval framework for multi-hop QA combining structured knowledge graphs and unstructured documents in semi-structured knowledge bases.
OMAC framework for optimizing multi-agent LLM systems, addressing handcrafted development approaches for collaborative agents in complex reasoning tasks.
LoVeC uses reinforcement learning to improve verbalized confidence in LLM generations, addressing hallucination detection without expensive sampling methods.
ReasonCache optimizes inference serving for large reasoning models by sharing KV cache to reduce memory overhead and improve throughput for concurrent requests.
Research on decomposing neural network representation spaces into interpretable subspaces using unsupervised learning methods for mechanistic interpretability.
Parameter-efficient architecture expansion using nested LoRA for multimodal LLMs to mitigate catastrophic forgetting during continual learning.
Research identifying performance pitfalls of KV cache compression in LLMs under realistic multi-instruction prompting scenarios.
Token-level knowledge transfer method enabling LoRA adaptation portability across different LLM backbones via contrastive learning.
Vision Expert Transformer distilling multiple foundation models for flexible robot learning via dynamic routing and feature selection.
Knowledge distillation method for LLMs using alpha-mixture assistant distribution to reduce computational costs while maintaining performance.
Research reconstructing visual stimuli from fMRI signals using latent space transformation and generative models.
Tutorial on cognitive biases in LLM-powered agentic AI for 6G autonomous networks using multimodal reasoning.
WAR-R1: explainable Web API recommendation system using semantic reasoning for mashup development.
Study of moral susceptibility and robustness in LLMs under persona role-play using Moral Foundations Questionnaire.
DTPQA: benchmark for evaluating Vision-Language Models on traffic scene perception with distance annotations.
Multi-objective optimization framework for Chinese short-form creative content generation with explanation-driven verification.
Multi-agent LLM systems for PyTorch inference optimization, outperforming traditional compilers on GPU tuning.
Research on curriculum-based LLM pretraining showing learning rate decay wastes high-quality training data.
Theoretical analysis of goal-conditioned reinforcement learning optimality gaps from optimal control perspective.
Theoretical analysis comparing exact versus approximate symmetry in ML models, showing approximate symmetry is computationally easier with empirical benefits.
LangPrecip framework incorporates meteorological text as semantic constraints in multimodal precipitation nowcasting to improve spatiotemporal forecasting.
PolyBench dataset and training approach teaches LLMs polymer design reasoning via domain-specific knowledge to overcome capability gaps in chemistry-related tasks.
OPT-ENGINE benchmark evaluates LLM capabilities in optimization modeling across OR problems, systematically scaling complexity from linear to mixed-integer programming.
L2R proposes low-rank and Lipschitz-controlled routing for Mixture-of-Experts models to improve expert specialization and routing discrimination in conditional computation.
CELM foundation model for end-to-end clinical EEG report generation from long-duration variable-length EEG recordings.
Lightweight patch-based representation learning method for time-series anomaly detection avoiding computational overhead of large models.
Reinforcement learning approach for LLM reasoning using human-inspired reward shaping with distinct exploration and consolidation stages.
Open-source framework and analysis quantifying fabricated citations in academic papers generated or assisted by LLMs.
Vision-language reasoning benchmark for remote sensing with 2,488 samples requiring complex reasoning beyond perception tasks.
Offline reinforcement learning method combining behavior cloning with actor-critic to address performance ceiling with suboptimal datasets.
Attention mechanism addressing representation collapse and attention sink phenomena through bounded confidence dynamics.
Transformer architecture for learning solution operators on complex geometries with parametric physical settings.
Improved circuit-tracing method ACC++ for identifying attention head mechanisms and interpretable circuits in language models.
Multi-agent LLM and vision framework for closed-loop robotic manipulation with environmental feedback.
GraphRAG-based framework for automated curation of clinical concept sets from medical text for NLP applications.
Privacy-preserving unlearning framework for LLMs enabling knowledge removal without sharing server parameters or forget sets.
Scalable reward modeling framework for robots using trajectory comparisons instead of absolute progress labels.
Benchmark for evaluating AI code generation models on complete web application development tasks with 964 browser-based workflows.
TERMINATOR learns optimal stopping points for chain-of-thought reasoning in LLMs to reduce computational waste from overthinking while maintaining answer quality.
Reinforcement learning methods for diffusion language models using entropy-guided step selection and stepwise advantage estimation without surrogate likelihoods.
M²RNN architecture with matrix-valued hidden states enabling non-linear RNNs for language modeling tasks requiring higher complexity than Transformer TC⁰ class.
Abduction-based debugging approach for LLM refinement on abstract reasoning tasks, formally re-checking transformations instead of outcome-level observation.
OneSearch-V2 generative retrieval framework using latent reasoning and self-distillation for improved complex query understanding in industrial search systems.
Systematic security analysis of OpenClaw AI agent framework, cataloging 470 vulnerabilities across architectural layers in LLM-based agent runtimes.
Studies post-training methods to make LLMs explicitly signal uncertainty in their responses, reducing confident yet incorrect outputs in real-world applications.
PinpointQA dataset and benchmark for evaluating multimodal LLMs on small object localization and spatial reasoning in indoor videos.
ECHO framework for speculative decoding in LLM inference with dynamic sparse gating, optimized for high-concurrency production serving scenarios.
RoboLab simulation benchmark for evaluating task generalist robotic policies with reduced domain overlap between training and evaluation to test true generalization.