BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
BioMedArena: open-source toolkit for building and evaluating biomedical research agents, reducing per-paper engineering overhead.
BioMedArena: open-source toolkit for building and evaluating biomedical research agents, reducing per-paper engineering overhead.
PAGE method for optimizing LoRA adapter placement in parameter-efficient fine-tuning of large language models.
Event-Causal RAG framework for long video reasoning in vision-language models addressing temporal coherence and causal inference.
Evaluation of zero-shot and few-shot LLMs for clinical action extraction from discharge notes using a two-stage framework.
Research on internal representations of social role granularity in LLMs using contrast-based latent directions and hidden state analysis.
VL-LCM: annotation-free evaluation framework for multimodal LLMs based on vision-language logical consistency metrics.
Dynamic Boundary Evaluation for adaptive LLM benchmarking that locates model-specific capability boundaries beyond fixed test sets.
Joint Consistency: test-time aggregation framework using energy minimization to aggregate multiple reasoning traces from LLMs.
Hygieia: multi-modal AI agent for rare disease diagnosis integrating phenotypic, genetic, and clinical data sources.
Safactory: infrastructure for scalable agent development covering evaluation, data management, and continuous improvement loops.
Data Language Models: foundation model class natively processing tabular data without preprocessing pipelines.
LLM-based taxonomy-agnostic PII detection in HTTP traffic with minimal labelled data.
Black-box method to estimate confidence in chain-of-thought reasoning using trajectory geometry and convergence analysis.
Framework for selecting optimal LLM controller classes for routing decisions based on input expressivity tradeoffs.
Distributional analysis of real vs synthetic pre-training data for tabular foundation models.
InciteResearch multi-agent framework for research ideation before questions are formed, automating tacit research friction.
Theoretical framework for agency under partial observability using bridge interfaces for sensing and actuation.
Execution lineage model for reproducible LLM agent workflows, preserving stable artifacts across tool use and refinement.
Analysis of risks from using AI agents to automate alignment research, including potential for misleading safety assessments.
Knowledge graphs improve LLM-based agentic workflows for SystemVerilog assertion synthesis in formal verification.
SCRuB framework for evaluating LLM reasoning about social concepts using rubric-based methodology.
PrefixGuard generates online failure-warning monitors for LLM agents from execution traces using induction and supervised learning.
Agentic Success Rate metric for evaluating LLM-based multi-agent payment workflows beyond task success.
Framework using graph kernels to analyze transformer circuits through activation patching for mechanistic interpretability.
ReasonSTL translates natural language to Signal Temporal Logic using tool-augmented learning for autonomous systems verification.
Introduces benchmark measuring instrumental convergence behaviors in LLM agents, testing tendencies to violate instructions for goal-relevant actions.
Uses Weisfeiler-Lehman graph analysis to examine co-occurrence structure in sparse autoencoder features for mechanistic interpretability.
Proposes process-based human-machine discrimination using cognitive science methods instead of output indistinguishability for autonomous agents.
Studies failure modes in RL-trained pricing agents where standard metrics hide poor market behavior, introducing trace diagnostics for diagnosis.
Framework for evaluating AI-induced diversity collapse in creative systems without requiring human-AI interaction data.
Deterministic adjoint matching framework for fine-tuning flow-based generative models via optimal control over velocity fields.
LLM agent system for coordinating multimodal neuroimaging analysis workflows including preprocessing, quality control, and statistical analysis.
Framework for LLM agents to learn and curate reusable skills from streaming tasks enabling self-evolution and continuous improvement.
Joint prompt optimization framework for LLM-based multi-agent systems addressing coordination of role-specific agent prompts.
Study of reinforcement learning for improving LLM long-horizon reasoning via controlled synthetic environment examining task difficulty and expressiveness.
Interactive workbench enabling mathematicians to leverage AI agents for exploratory research including literature search, computation, and theorem proving.
Theoretical analysis questioning whether flat minima in loss landscapes causally explain generalization or are artifacts of parameterization.
Review of LLM applications in quantitative finance for stock price forecasting, sentiment analysis, and multi-agent trading systems.
Adaptation of Mamba architecture for medical time series classification with improved long-range dependency capture.
Physics-informed neural networks with learnable loss balancing for scientific machine learning under data scarcity.
Framework characterizing model multiplicity in chaotic systems using horizon-constrained Rashomon sets.
Sparse prefix caching optimization for LLM serving that exploits state-space model structure for improved latency.
Theoretical framework for steering intermediate representations in generative models via affine transformations for alignment and safety.
Token-Selective Attention mechanism for adaptive computation depth in transformers, reducing unnecessary layer computation per token.
Theoretical analysis of sparse autoencoders for disentangling feature superposition in transformers and compositional steering mechanisms.
Load balancing optimization for efficient multimodal MoE LLM inference addressing information heterogeneity.
Framework converting outcome-level supervision into process supervision for reinforcement learning in reasoning tasks.
Online reweighting approach for data curation in LLM training that outperforms offline selection methods.
Evolutionary fine-tuning of quantized deep learning models for compression on IoT and mobile devices.
Direct corpus interaction approach for agentic search beyond semantic similarity, enabling multi-step reasoning.