ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space
Proposes ABC method for generating continuous stochastic processes conditioned on partial observations using non-Markovian diffusion bridges.
Proposes ABC method for generating continuous stochastic processes conditioned on partial observations using non-Markovian diffusion bridges.
Research on sampler-robust optimization for stochastic pipelines using generative models. Addresses sampler misspecification and Monte Carlo simulation errors.
Combines hyperbolic geometry and diffusion models to improve graph few-shot learning performance.
Systematic review of security attack and defense strategies for LLM-based autonomous agent frameworks with OpenClaw case study.
Proposes causally-motivated inference-time intervention to mitigate multiple biases in LLM reward models beyond response length.
Framework for identifying when to ask humans versus AI agents during information seeking in hybrid environments.
Parallel corpus dataset for summarizing and interpreting privacy policies using NLP methods.
Evaluates generalization capabilities of transformer models for program synthesis using controlled domain-specific grammar experiments.
Method for hierarchical alignment between medical images and radiology reports using deep learning.
ZAYAN self-supervised contrastive framework for tabular remote sensing data, performing feature-level rather than sample-level contrastive learning without anchor selection.
Cognitive Digital Shadows dataset of 190k LLM responses across 19 models debating societal issues while shadowing human personas and demographics for discourse analysis.
HAVEN hybrid verification engine leveraging LLMs to generate hardware testbenches while correcting common HDL mistakes through automated verification feedback loops.
ANCORA framework shifting from answer-generation to question-generation, using self-play between proposer and solver agents for verifiable reasoning without human supervision.
Analysis of stability-plasticity tradeoff in continual learning, examining when modular structure benefits or harms task transfer and interference based on dimensionality.
Study identifying hubness vulnerabilities in CLIP and cross-modal encoders, showing single hub text embeddings can degrade retrieval performance across unrelated examples.
VibroML open-source Python toolkit for automated structural remediation of crystalline materials using machine-learned potentials and phonon analysis.
AgentEconomist is an agentic system translating economic intuitions into executable computational experiments, grounded in 13k+ academic papers with multi-agent agentic workflows.
Position-aware speculative decoding acceleration technique for LLM-based generative list-wise recommendation systems, optimizing inference latency while preserving target distribution.
Method for integrating domain knowledge into neural networks via Differentiable Knowledge Units that discover and learn symbolic rules for improved vision task generalization.
LLM-based poetry generation system for Arabic and its dialects with instruction-guided generation, extending beyond prior analysis-focused tasks to creative composition.
RuC benchmark for evaluating LLMs on hardware description language (HDL) rule completion tasks, addressing gaps in code completion evaluation for RTL development.
Framework addressing behavioral drift from silent LLM provider updates in production systems, proposing versioning and testing approaches for LLM supply chain governance.
Empirical study comparing Google Search, Gemini, and AI Overviews on 11,500 queries, analyzing how generative AI disrupts traditional web search results and source presentation.
CastFlow presents an LLM-based agentic system with role-specialized workflows for time series forecasting, moving beyond single-pass static generation to multi-round iterative refinement.
NeocorRAG proposes Recall Conversion Rate metric to measure how retrieval improvements in RAG systems translate to downstream reasoning accuracy, addressing gaps between retrieval and reasoning performance.
Evaluation of small language models (EuroLLM, Aya, Gemma) on emotion preservation in machine translation backtranslation tasks.
Survey of AI-assisted peer review covering LLM-based generation, agent systems, RL methods, and automation across review pipeline stages.
Multimodal LLM framework for reliable circuit-diagram-to-Verilog code generation with grounding for hardware design accuracy.
TransVLM vision-language model detects shot transitions in video as continuous temporal segments rather than cut points.
ITS-Mina proposes MLP-based architecture with Harris Hawks optimization for multivariate time series forecasting.
Framework using clinician overrides as preference signals for clinical AI training via RLHF-style learning from expert disagreement.
LLM-based approach to Design Structure Matrix modularization for system engineering, using language models for combinatorial optimization.
Template Constrained Decoding improves LLM-based Text-to-SQL generation for reliable database querying with complex schemas.
MIFair framework addresses bias and fairness in ML using mutual information, handling intersectionality and multiclass settings.
Research examining pre-deployment factors determining whether AI systems are built or abandoned, focusing on responsible AI development decisions.
Research on high-quality data filtering for German language model training, comparing strict quality filtering vs. diverse large-scale approaches.
TopBench benchmark for implicit prediction and reasoning in tabular QA, testing LLMs on predictive queries beyond simple retrieval.
Neuro-symbolic framework combining first-order logic, causal models, and RL for rule synthesis in safety-critical domains.
DEFault++ hierarchical fault diagnosis tool for detecting and categorizing failures in Transformer architecture components.
Research examining whether sparse autoencoders capture concept manifolds versus independent linear directions in neural networks.
PRISM framework for multimodal model pre-alignment using black-box on-policy distillation before supervised fine-tuning and RLHF.
AdvDMD combines adversarial rewards with distribution matching distillation to improve few-step diffusion model generation quality.
Method detecting multi-turn prompt injection attacks via LLM activation patterns and adversarial restlessness signatures in residual streams.
Crab runtime for checkpoint/restore in agent sandboxes enabling fault tolerance, spot execution, and safe rollback for autonomous agents.
Claw-Eval-Live benchmark for evaluating LLM agents on evolving real-world workflows with live signal layer and execution verification.
OpenAI o1 system documentation describing reinforcement learning training, chain-of-thought reasoning, and safety alignment.
Research on adversarial attacks against LLMs through deductive and inductive illusions, tested with chain-of-thought reasoning.
Framework for tracking provenance and chronological history of multi-agent collaborative content generation and revisions.
LLM-based agent system for optimizing chip design by learning from thousands of parameters and complex engineering workflows.
Programmatic platform for policy-grounded safety evaluation of LLMs using adversarial prompts and AI-based rater validation.