Learning to Present: Inverse Specification Rewards for Agentic Slide Generation
RL environment where LLM agents learn to generate professional presentations through research, planning, and tool use with multi-component reward system.
RL environment where LLM agents learn to generate professional presentations through research, planning, and tool use with multi-component reward system.
Method for training LLM agents to leverage rich environment feedback through reflective experience and post-training, improving long-horizon planning.
Benchmark evaluating audio-visual social interactivity capabilities of omni-modal LLMs in dynamic dialogue settings.
RL framework using Soft Actor-Critic to learn adaptive ray sampling policies for efficient neural radiance field rendering.
Multimodal AI search framework combining vector search, hybrid retrieval, and reasoning for pharmaceutical data across text, images, audio, and video.
Evaluation of VLMs (GPT-4V, Gemini, Claude, Llava) for navigation assistance tasks for people with vision impairments.
Framework extending RLHF using multi-dimensional rubric-based rewards instead of scalar signals for RL training.
Inference-time governance approach for LLMs using adaptive prompt routing to enable social alignment without retraining.
Study on recursive language models with self-reflective program search for long-context handling, addressing information extraction challenges.
Analysis of Gini Index role in prompt-based classification for detecting and optimizing class accuracy disparities in long-tailed datasets.
Defense mechanism against steganographic collusion in multi-agent reinforcement learning using dynamic representational circuit breaking.
Model rectification framework using attribution-guided rank-one editing to fix unreliable neural network behaviors on corrupted samples.
Open-source pipeline extending single-agent AI orthodontic treatment planning to dual-agent framework with improved tooth segmentation and landmarks.
AI agent system for hardware design reviews using LLMs to verify semantic correctness of component connections against datasheets.
Framework for LLM application release management using automated self-testing with evidence-based quality gates across five dimensions.
Analysis of transformer training dynamics using Spectral Edge Dynamics to measure coherent optimization directions versus stochastic noise.
Context-aware safety framework for personalized text-to-image models that prevents misuse without broad concept erasure.
Analysis of multi-turn safety failures in LLMs through state-space perspective, showing structured contextual evolution enables jailbreaks.
Token compression method for omnimodal LLMs using dynamic audio-driven semantic chunking to reduce inference costs for audio-visual processing.
Study on engineering challenges in LLM-based multi-agent systems, addressing context pressure, coordination errors, and system drift at scale.
Defense framework against backdoor attacks in LLMs using trigger generation and inversion to locate and mitigate malicious triggers.
Research on using inference time as a proxy to estimate LLM energy consumption, addressing opacity in API-based model access and environmental impact.
SEMAG: self-evolutionary multi-agent code generation framework that decomposes programming tasks into planning, coding, debugging stages with adaptive workflow selection.
Uncertainty-guided multi-expert framework for imbalanced sequence learning addressing poor expert specialization and prediction conflicts in long-tailed data.
Retrieval-augmented generation framework using GPT-4 to accelerate CO2 reduction catalyst discovery by exploring chemical spaces and interpreting results.
Method bridging learned embeddings and handcrafted features in event sequences for financial systems, addressing interpretability and latency constraints in production ML.
Large-scale competition analysis revealing LLM agents' vulnerability to indirect prompt injection attacks through adversarial instructions in external content sources.
Meta-TTRL: metacognitive test-time reinforcement learning framework for unified multimodal models enabling knowledge accumulation across similar prompts in text-to-image generation.
MiroThinker-1.7 and H1: research agents with enhanced verification and multi-step reasoning via structured planning and contextual reasoning for long-horizon tasks.
ClawWorm: first documented self-propagating attack across LLM agent ecosystems, demonstrating security vulnerabilities in OpenClaw platform with 40,000+ active instances.
Technique improving pretrained diffusion/flow-matching robot policies by replacing sampled noise with optimized constant vectors for better downstream reward performance.
Simulation Distillation method enabling sim-to-real transfer in robotics by pretraining world models in simulation for rapid real-world adaptation with low data.
CorrectionPlanner: autonomous driving planner using reinforcement learning with explicit self-correction mechanism in propose-evaluate-correct loop for unsafe action handling.
Evaluation of how LLMs and tokenizers handle Arabic root-pattern morphology, testing whether models capture genuine morphological structure or rely on surface memorization.
Differentiable framework for computing geodesics on 3D meshes with parallelization support to improve machine learning on non-Euclidean geometric domains.
OMNIFLOW: multimodal agent combining LLMs with physics-grounded reasoning to handle spatiotemporal PDE dynamics without domain-specific fine-tuning, reducing non-physical hallucinations.
Security framework for LLM-based multi-agent systems addressing manipulation risks from malicious agents in interactive agentic networks through communication channel exploitation.
Large-scale empirical study demonstrating prediction-equivalent models produce substantially different feature attributions across 24 datasets, challenging explainability assumptions.
Study showing LLM stability across repeated runs does not guarantee agreement with statistical ground truth in data-constrained scientific decision-making workflows.
Informationally Compressive Anonymization (ICA) method for privacy-preserving ML that protects sensitive data without the performance degradation of differential privacy or homomorphic encryption.
arXiv: Design principles for XAI interfaces enabling scientists to probe and interpret LLM behavior in reading and research workflows.
arXiv: Counteractive RL framework addressing exponential state space complexity for efficient deep reinforcement learning.
arXiv: LLM ensemble approaches for word sense plausibility rating in SemEval-2026 using zero-shot and Chain-of-Thought prompting.
arXiv: Framework for Internet of Physical AI Agents addressing interoperability, security, and sustainability in IoT environments.
arXiv: Privacy-preserving federated learning for Alzheimer's classification using 3D MRI with site-aware techniques.
arXiv: Practical guide to AI-assisted research in mathematics and ML, covering productive tool use and responsible guardrails.
arXiv: Analysis of 10,469 experiments by Claude Opus and Gemini agents across 108k design space cells for ML architecture search.
arXiv: VIBEPASS empirically evaluates LLM self-diagnosis and repair capabilities for autonomous software engineering.
arXiv: Benchmarking causal discovery algorithms on synthetic healthcare data for fairness and utility evaluation.
arXiv: LLM-guided neural architecture search for multimodal time-series classification under data-locality constraints for healthcare.