Generative Evolutionary Meta-Solver (GEMS): Scalable Surrogate-Free Multi-Agent Reinforcement Learning
Surrogate-free multi-agent reinforcement learning framework using generative models instead of explicit policy populations.
Surrogate-free multi-agent reinforcement learning framework using generative models instead of explicit policy populations.
Transformer architecture using cross-state transition attention for robust robotic manipulation from demonstrations.
Prompting protocol combining objection-raising and revision mechanisms to improve LLM reasoning and self-correction.
Multi-turn red-teaming approach using tree-based dialogue and reinforcement learning for discovering LLM vulnerabilities.
Hardware-software co-design framework for efficient multimodal model inference on battery-powered edge devices.
Membership inference attacks on LLM tokenizers as privacy attack surface distinct from model attacks.
Backdoor attack on vision-language-action models demonstrating action-level behavioral manipulation vulnerabilities.
World model and MPC framework for humanoid robot contact planning combining learned representations with sampling-based control.
Open-source corpus and tools for training fully open multimodal LLMs with improved data quality and reasoning.
Study on unintended reasoning behaviors in reinforcement-learning-trained LLMs and chain-of-thought monitoring.
Continual learning method for audio-visual segmentation addressing modality entanglement in sequential tasks.
Framework enabling LLMs to perform tabular prediction via structural priors and reasoning-focused optimization.
Evaluates driving world models as synthetic data generators for autonomous vehicle perception tasks.
Navigation system using 3D Gaussian Splatting memory for multi-modal visual goal navigation in robotics.
SwiftEmbed: production text embedding system achieving 1.12ms latency and 50k req/s using static token lookup in Rust.
Research on vectorized online POMDP planning for autonomous robot decision-making under partial observability with parallelization.
Research on detecting AI-generated images via diffusion model snap-back reconstruction forensics. Addresses Stable Diffusion and DALL-E detection.
Comparative study of interpretable fuzzy reasoning vs deep learning for motor-imagery EEG classification in brain-computer interfaces.
Research paper on federated learning of mixture-of-experts models for mobile edge computing and resource-constrained devices.
FATE benchmark series for formal algebra theorem proving at multiple difficulty levels. Evaluates LLM capabilities on mathematical reasoning beyond contest problems.
Detection method for AI-generated images using contextual anomaly estimation in masked autoencoders. Extends DetectGPT approach from text to vision domain.
HatePrototypes: Interpretable representations for hate speech detection covering implicit and explicit hate. Addresses content moderation with transferable embeddings.
UnfoldLDM combines deep unfolding networks with latent diffusion models for blind image restoration. Model-based interpretable approach to image processing.
Probabilistic certification framework improving SmoothLLM defense against LLM jailbreaking attacks. Addresses robustness guarantees with realistic assumptions.
Yo'City: Agentic framework using self-critic expansion for personalized, boundless 3D city generation. Demonstrates AI agent reasoning in creative generation tasks.
Automated pipeline for generating multi-turn conversational jailbreak attacks against LLMs using psychological principles like FITD without manual dataset creation.
Contrastive learning approach for adapting foundation models to domain-specific tasks in Earth observation without full retraining.
AltNet addresses plasticity loss in RL-trained neural networks via parameter reset strategies. Research on continual learning for RL agents.
arXiv paper on evaluating agentic systems via process-centric analysis of trajectories and reasoning patterns rather than outcomes alone. Foundational agent analysis framework.
SALVE framework for neural network interpretability and control using sparse autoencoders. Mechanistic interpretability research not focused on LLMs or agents.
LaMer: Meta-RL framework enabling LLM agents to actively explore and learn from trial-and-error in multi-turn tasks. Research on agent training methodology.
arXiv paper analyzing cost trade-offs between reasoning and non-reasoning LLMs for Text-to-SQL tasks on cloud platforms. Empirical efficiency comparison.
DrivingGen benchmark for generative video world models in autonomous driving. Research on agent simulation and synthetic data generation.
NC-Bench: arXiv benchmark evaluating LLM conversational competence on form/structure vs content. Research paper on LLM evaluation methodology.
Audit of LAION-Aesthetics Predictor studying whose aesthetic values are embedded in visual generative AI training datasets.
Benchmark with 1,800 code completion instances across 6 languages derived from real developer telemetry; avoids contamination, enables detailed diagnostics.
Training-free caching framework for Flow Matching inference using average-velocity perspective and Jacobian-vector products for acceleration.
Open-source cybersecurity LLM trained on 11.8B tokens of curated domain data; supports diverse security workflows while protecting sensitive data.
Data Shapley attribution method for adaptive optimizers like Adam, extending in-run attribution beyond SGD's linear structure.
Multi-label classification study using Schwartz value hierarchies for sentence-level human value detection on sparse, imbalanced datasets.
Reward shaping method for LLM reasoning via reinforcement learning, addressing entropy collapse and exploration challenges in verification-based training.
Study of recurring vulnerabilities in LLM-generated code; introduces FSTab for black-box attacks predicting backend security issues from frontend patterns.
Semantic search system over 9M mathematical theorems using embeddings to retrieve specific results for mathematicians and theorem-proving agents.
LLM-driven recommendation system using multimodal motivation modeling to improve content preference prediction by incorporating review text and heterogeneous data.
Diffusion-guided pretraining for brain graph foundation models, using learnable augmentation for connectome data instead of random dropping.
CoCoA decoder mitigates LLM hallucinations by detecting representational instability across layers, requiring no training.
SToRM token reduction technique optimizes multimodal LLMs for end-to-end autonomous driving with natural language interaction.
Uses agent guidance from learned policies to accelerate robotic RL, reducing sample inefficiency without 1:1 human supervision.
TrasMuon optimizer improves Muon-style methods by adding trust-region adaptive scaling for robust gradient updates.
Variational flow-matching framework for simulation-based inference with structured domain constraints on posteriors.