MARVEL: Multi-Agent RTL Vulnerability Extraction using Large Language Models
Multi-agent system using LLMs to detect vulnerabilities in hardware RTL design specifications.
Multi-agent system using LLMs to detect vulnerabilities in hardware RTL design specifications.
Multimodal LLM combining vision, audio, and sensor data for embodied AI agents in smart homes.
Benchmark for evaluating multimodal LLM capabilities in humanities and social sciences requiring interdisciplinary reasoning.
Comprehensive benchmark for evaluating multimodal LLMs on front-end code generation from visual designs using modern frameworks.
Framework addressing dependency, asynchrony, and missing values in multivariate time series forecasting from real-world data.
LLM-based framework for evaluating children's language function through phonetic transcription and automated speech assessment.
Characterization and comparison of State Space Models and hybrid architectures versus Transformers for long-context processing on edge devices.
Monte Carlo tree diffusion with multiple experts for protein sequence design combining masked diffusion models with tree search.
Framework using LLMs as spatio-temporal predictors with hierarchical temporal tokenization for human mobility and trajectory prediction.
Reinforcement learning fine-tuning approach using polychromic objectives to prevent policy collapse and maintain behavioral diversity.
Learnable dynamic routing mechanism for mixture-of-experts with LoRA adapters enabling efficient LLM task adaptation without fixed expert assignment.
Policy optimization method for text-to-image models addressing credit assignment instability in reinforcement learning fine-tuning.
Evaluation of Vision-Language-Action model robustness against multi-modal perturbations across 17 adversarial conditions.
LLM-based agent framework for recommendation systems that leverages commonsense reasoning to capture item relationships and user intent.
Study on how LLM-generated rationales influence human plausibility judgments in commonsense reasoning tasks using 3,000 human and 13,600 LLM judgments.
Latent-Augmented Discrete Diffusion model improving fast language generation by modeling cross-token dependencies.
Method for scalable AI oversight via partitioned human supervision across multiple domain experts for complex multi-domain tasks.
Security evaluation framework for LLM backbone models in AI agents, addressing vulnerabilities unique to agent architectures.
Survey examining terminology, definitions, and taxonomy of data agents—autonomous systems orchestrating data and AI for complex data tasks.
LLM-based search agents trained on synthetic entity-centric data using improved reward mechanisms to capture informative near-miss samples.
OckBench: Benchmark measuring LLM reasoning efficiency via token usage, revealing up to 5x differences in token length across models.
Data-efficient fine-tuning strategy for adding controllable parameters to text-to-video diffusion models using synthetic data.
Refusal Steering: Inference-time method for fine-grained control over LLM refusal behavior on sensitive topics without retraining.
HiGR: Generative slate recommendation system using hierarchical planning and multi-objective preference alignment for ranked lists.
CogFlow: Multimodal LLM system for visual math problem solving improving visual perception integration and reasoning.
Comprehensive empirical study evaluating factors affecting safety alignment in LLMs and LRMs across 32 recent models.
Fast-ThinkAct: Efficient Vision-Language-Action framework reducing inference latency through verbalizable latent planning.
CLiMB: Domain-informed clustering framework for novelty detection in scientific discovery with application to galactic archaeology.
Molmo2: Open-weight vision-language model with video understanding, grounding, and disclosed training data and recipe.
FROST: Attention-aware pruning method for efficient LLM reasoning by identifying and removing reasoning outliers while preserving capacity.
Persona Brainstorm Audit method for detecting bias and fairness issues in open-ended creative outputs from LLMs.
Study showing SGD with sparsity outperforms Adam for RL from verifiable reward in LLM training, challenging standard optimization practices.
AceGRPO: Reinforcement learning agent for autonomous ML engineering using adaptive curriculum and group relative policy optimization to overcome parameter freezing.
Study on paraphrase generation and detection as mechanisms for language understanding and modeling in neural networks.
UI-Venus-1.5: GUI agent with 2B, 8B, and 30B-A3B variants for automating digital environment interactions with broad generality and strong task performance.
VESPO: reinforcement learning method for training LLMs with improved stability through soft policy optimization and importance sampling to address policy divergence.
KBVQ-MoE: compression technique for Mixture of Experts LLMs using vector quantization and SVD to reduce parameter size and memory for resource-constrained deployment.
Sim2Radar uses VLM-guided scene reconstruction to synthesize radar training data from RGB images, addressing radar dataset scarcity.
Identifies and analyzes 'silent inconsistency' problem in data-parallel LLM fine-tuning where gradient synchronization doesn't ensure worker-level optimization alignment.
ST-EVO framework for self-evolving LLM-powered multi-agent systems that dynamically construct task-adaptive communication topologies instead of predefined structures.
Studies mechanical basis of capability emergence in neural networks across scales 405K-85M parameters, finding scale-invariant representation collapse and top-down reorganization.
Proposes AI-CARE metric incorporating carbon emissions and energy consumption alongside standard performance metrics for ML model evaluation.
Analyzes quality issues in AI safety datasets, finding they rely on superficial 'triggering cues' rather than genuine adversarial patterns.
Randomized trial showing AI-generated feedback suggestions via FeedbackWriter improve student revisions when reviewed by human TAs in economics course.
Proposes symbolic alternative to GNN message-passing for more interpretable and expressive graph learning in high-stakes domains.
MASPO algorithm improves LLM reasoning through reinforcement learning with verifiable rewards, addressing gradient utilization and probability mass issues in existing RLVR methods.
User study comparing chatbots, games, and essays for persuasive learning on sustainability topics with identical content.
Uses LLM-assisted reasoning to map 2D engineering drawing annotations to 3D CAD features for manufacturing automation and process planning.
Benchmark study (SP-ABCBench) evaluating whether LLM agents can simulate human security and privacy attitudes and behaviors for risk forecasting.
Examines integration of AI into science education materials, covering personalization, adaptive instruction, and accessibility in K-12 learning contexts.