NeuDiff Agent: A Governed AI Workflow for Single-Crystal Neutron Crystallography
NeuDiff Agent LLM-based workflow for automated analysis and reporting in neutron crystallography at Spallation Neutron Source.
NeuDiff Agent LLM-based workflow for automated analysis and reporting in neutron crystallography at Spallation Neutron Source.
Node Learning decentralized paradigm for edge AI where intelligence resides at individual nodes without centralized servers.
IndicJR judge-free benchmark of jailbreak robustness across 12 Indic/South Asian languages covering 45,216 adversarial prompts.
GUI-Owl-1.5 native GUI agent model in multiple sizes supporting desktop, mobile, browser with state-of-the-art results on 20+ automation benchmarks.
OpenSage first agent development kit with self-programming capability for automatically designing agent topology, tools, and memory components.
AgentLAB benchmark for evaluating LLM agent vulnerabilities to adaptive long-horizon attacks in complex multi-turn environments.
LLM-WikiRace benchmark evaluates planning, reasoning, and world knowledge by requiring models to navigate Wikipedia hyperlinks from source to target page.
Study showing fine-tuning vision-language agents on narrow tasks causes emergent misalignment that generalizes across unrelated domains and modalities.
DeepContext stateful monitoring framework for detecting adversarial intent drift across multi-turn LLM dialogues, addressing safety gaps in sequential interactions.
SourceBench evaluates quality of web sources cited by LLMs across 100 queries using eight-metric framework beyond correctness.
GAP benchmark reveals that text-level safety alignment in LLM agents doesn't transfer to tool-call safety, measuring real-world action harms.
LLM4Cov framework for offline agent learning applied to high-coverage hardware testbench generation using non-differentiable execution feedback.
Phantom: automated agent hijacking attack on LLM agents via structural template injection, addressing OWASP-highlighted threat with improved transferability.
Theoretical analysis of fundamental limits in black-box safety evaluation of AI systems, showing latent context-conditioned policies create evaluation gaps.
Conv-FinRe benchmark for stock recommendation that evaluates utility-grounded decisions rather than behavioral imitation in conversational finance advisory.
Sonar-TS neuro-symbolic framework for natural language querying time series databases, handling morphological intents and ultra-long histories.
M2F agentic framework for end-to-end project-scale autoformalization of mathematics in Lean, managing cross-file dependencies and imports.
AI agent for Microsoft Dynamics 365 Sales querying live CRM data, reasoning over schemas, and producing decision-ready insights with benchmarking.
Mixture-of-Experts architecture for RL policy networks in LLM agents, addressing simplicity bias by allocating capacity across task complexity.
Instruction-Tool Retrieval (ITR) RAG variant dynamically retrieves minimal system prompts and necessary tool subsets per step for efficient agentic LLMs.
Multi-agent computer-use framework with intent-aligned plan memory to stabilize long-horizon execution and reduce error accumulation.
Framework evaluating reasoning faithfulness in large reasoning models through counterfactual intervention on stance consistency and causal influence.
Multi-agent RL method retaining multiple high-value actions via sub-value functions to adapt to shifting value functions.
Predictive Batch Scheduling uses lightweight online predictor to prioritize high-loss samples, accelerating LLM training convergence.
Empirical study analyzing pull request characteristics from five AI coding agents and human reviewer responses using AIDev dataset.
6G wireless systems architecture using intent-driven autonomous agents for multi-dimensional objectives and evolving requirements.
Owen-value based method extending SHAP for hierarchical feature attribution in vision tasks with spatial/semantic dependencies.
Knowledge graphs capturing educational concept dependencies and prerequisites for personalized learning at scale.
Philosophical examination of generative AI's epistemic character and implications for knowledge production in science, education, and institutions.
Texo: minimalist 20M parameter formula recognition model achieving state-of-the-art performance with 80% size reduction through distillation and transfer learning.
Framework for constructing symbolic causal world models online by integrating continuous model learning with meta-interpretive learning in agent decision loops.
Methodological experiment using AI agents in collaborative research workflows for humanities and social sciences, analyzing Taiwan Claude.ai usage data.
Framework for predicting consistent individual-specific human behavior in high-stakes environments by combining LLMs with psychological trait modeling.
Mechanistic interpretability study using linear probing and Bloom's Taxonomy to analyze cognitive complexity in LLM internal neural representations.
Framework for detecting and quantifying temporal data contamination in LLM backtesting to validate whether models leak post-cutoff training knowledge.
Web Verbs framework providing typed abstractions for reliable task composition on agentic web, enabling LLM-based web agents beyond low-level primitives.
Case study training 1.36B-parameter scientific language model from raw arXiv LaTeX sources, documenting end-to-end process for domain-specialized LM development.
MedClarify: LLM-based AI agent for medical diagnosis that iteratively asks follow-up questions to resolve diagnostic uncertainty through differential reasoning.
Method for disentangling task vectors in foundation models using Kronecker-factored approximate curvature without external task data.
Graph-based visual inference approach for complex image retrieval queries involving relationships, compositions, and precise constraints.
Privacy-by-Design framework for LLM-based applications targeting children, addressing implementation gaps in privacy regulation compliance.
Benchmarking framework for optimizing AI models on ARM Cortex embedded processors, measuring energy efficiency, accuracy, and resource utilization.
arXiv paper on applying LLMs to telecom domain using dynamic knowledge graphs and retrieval-augmented generation to reduce hallucinations and improve accuracy.
Evaluation framework for chain-of-thought reasoning quality using reusability and verifiability metrics in multi-agent IR pipelines.
KLong open-source LLM agent framework for extremely long-horizon tasks using trajectory-splitting SFT and progressive RL training.
ODESteer unified ODE-based framework for LLM alignment via activation steering with multi-step guidance.
Federated learning ensemble combining SWIN Transformer and CNN for lung disease diagnosis from medical imaging.
AI Gamestore platform for evaluating machine general intelligence using open-ended human games and dynamic benchmarks.
MolHIT hierarchical discrete diffusion model for molecular graph generation improving chemical validity for drug discovery.
AutoNumerics multi-agent framework autonomously designs, implements, and verifies numerical PDE solvers using AI.