RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
RadAnnotate uses LLMs with retrieval augmentation and selective automation for efficient radiology report annotation.
RadAnnotate uses LLMs with retrieval augmentation and selective automation for efficient radiology report annotation.
FormulaCode benchmark evaluates LLM coding agents on repository-level codebase optimization with realistic multi-objective constraints.
Probing-based analysis of moral reasoning trajectories in LLMs across six models showing systematic multi-framework deliberation.
Critic-free RL approach for cross-user activity recognition from wearable sensors with temporal feature generation.
Framework adapts vision-language models as online reward generators for robotic reinforcement learning policy refinement.
Survey of resource consumption threats in LLMs including excessive generation, covering efficiency challenges for providers and users.
HEAR framework extends vision-language-action models to incorporate real-time sound for robotic manipulation tasks.
RecBundle proposes geometric framework for recommender systems addressing information cocoons through topological representation learning.
Inference-time repair layer for retrieval-grounded QA using answer-conditioned counterevidence retrieval to fix commitment errors.
Parallel in-context learning method reducing latency in vision-language models by decoupling demonstration processing from query encoding.
LLM serving system optimizing agentic workflows by handling cross-call dependencies and redundancy from speculative execution.
Data curation method for calibration in LLM compression via frequency-based selection for pruning and quantization.
Local-first multi-agent architecture for automated repository code review with LangGraph orchestration and structured analysis.
Automated skill distillation and adaptation method for financial reasoning in LLMs without fine-tuning.
Reference-free evaluation framework for pathology vision-language models to detect hallucinations without ground truth.
Benchmark for repository-level code understanding with executable environments, enabling agentic code automation tasks.
Benchmark comparing generative augmentation strategies (GANs, diffusion) for bias correction in imbalanced classification under low-data conditions.
Constrained RL method for enforcing hierarchical instruction priority in LLMs via system prompt compliance.
Transformer architecture for 4D point cloud video understanding with temporal scale invariance.
RL method preserving diversity in LLM reasoning via dynamic Jensen-Shannon replay to improve sample efficiency and avoid mode collapse.
Open-source reproduction of Corrective RAG replacing proprietary components with Wikipedia API and open models for improved reproducibility.
Local-first long-term memory system for AI assistants with vector and keyword retrieval, implemented in Rust for conversational agents.
Benchmark and method for evaluating 360° image perception in multimodal LLMs, addressing geometric distortion and spatial reasoning challenges.
Domain adversarial training approach for robust AI-generated audio quality assessment without spurious correlations.
Scoping review of AI-driven digital mental health interventions including GenAI and HCAI across screening, support, and monitoring.
CoMAI multi-agent framework with task decomposition for robust and fair interview evaluation using coordinated LLM agents.
Technical review and taxonomy of 13 generative systems for quantum circuit and quantum code generation including agentic approaches.
Visual prompt discovery method to diagnose and mitigate LVLM perception failures through semantic exploration.
Vision-language process reward models with explicit visual premise verification for reliable step scoring in reasoning.
Genetic programming with surrogate models for dynamic multi-mode project scheduling with simulation-based optimization.
VisBrowse-Bench benchmark for evaluating visual-native search in multimodal browsing agents using MLLMs.
End-to-end framework using Speech LLMs for spoken question answering with attention-guided evidence grounding.
Human-centered architecture for integrating LLM-based cognitive assistants into manufacturing quality management systems.
Security research on sentiment steering attacks targeting RAG-enabled large language models and LLM robustness.
Machine learning pipelines for radio astronomy data processing with explainability focus on automating configuration.
YOLO-based deep learning for automated wasp identification with explainable AI integration for taxonomic classification.
DynamicGate MLP conditional computation framework using learned structural dropout and input-dependent gating for efficiency.
FederatedFactory zero-dependency framework for federated learning in non-IID scenarios using generative one-shot learning.
Physics-guided diffusion framework for full-waveform inversion combining score-based generative models with wave-equation simulations.
Fanar 2.0 Arabic generative AI platform built on 256 H100 GPUs at QCRI with sovereign infrastructure and data pipelines.
Study identifying flaws in LLM benchmarks for Icelandic, highlighting issues with synthetic and machine-translated evaluation data.
PlotTwist creative plot generation framework using small language models with specialized training for narrative coherence.
Method for adding persistent memory to frozen encoder-decoder LLMs via trainable adapters in continuous latent space.
IndexRAG approach for multi-hop question answering that performs cross-document reasoning at indexing time using bridge entities.
SF-Mamba state space model for vision addressing non-causal patch interactions with improved computational efficiency.
SlideFormer system for fine-tuning large language models on single GPU via asynchronous engine and sliding window approach.
EngGPT2-16B open Italian LLM trained on 2.5T tokens, efficient inference with performance comparable to larger models.
Multi-agent reinforcement learning approach for managing delayed channel state information in multi-satellite communication systems.
Unlearning method for one-step generative models using unbalanced optimal transport for safer image generation.
LenghuSky-8 millisecond-resolution network dataset for time series foundation models with high-frequency data.