AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
AI4EOSC federated cloud platform for AI in scientific research with reproducible ML lifecycle management across distributed e-Infrastructures.
AI4EOSC federated cloud platform for AI in scientific research with reproducible ML lifecycle management across distributed e-Infrastructures.
SentGraph hierarchical sentence graph enables multi-hop retrieval-augmented generation by constructing coherent evidence chains from multiple documents for complex QA.
Audits political alignment in 26 LLMs using psychometric inventories and news bias labeling across prompt variants to evaluate behavioral bias.
Vision-language reasoning for urban socio-semantic segmentation from satellite imagery, distinguishing socially-defined categories beyond physical attributes.
Ablates RLVR pipeline components for code verifiers: intermediate thinking traces, negative samples, and on-policy training to reduce adoption cost barriers.
CFM language-aligned concept foundation model decomposes vision model representations into human-interpretable concepts with spatial grounding for diverse downstream tasks.
LREAD rubric-based framework with three-phase expert study calibrates human detection of LLM-generated Korean text to improve attribution accuracy beyond surface-level assessment.
Time-annealed perturbation sampling for diffusion language models enables diverse generation across semantic and reasoning paths by controlling temporal denoising.
Bauplan code-first lakehouse with data contracts, versioning, and transactional pipelines for concurrent AI/analytics workloads supporting both human and agent workflows.
Predicts LLM success from pre-generation internal activations using linear probes to enable efficient inference routing on math and coding tasks.
Benchmark evaluating multimodal LLMs' ability to understand pedagogical reasoning and science instruction in K-12 classroom videos with model-based explanations.
Descent-guided policy gradient method addresses cross-agent noise scaling in multi-agent reinforcement learning, improving sample complexity from O(N/ε) to sublinear bounds.
KEEP system optimizes KV-cache memory management for memory-augmented LLMs in embodied planning, reducing prompt length and prefill latency for long-horizon tasks.
Experimental study of sycophancy in LLMs—tendency to favor user-affirming over critical responses—with controlled interventions to identify and prevent the alignment failure.
Evaluates 5 open-source small LLMs (Gemma, Phi, Llama, Mistral, Meditron) for clinical QA reliability and prompt sensitivity in low-resource healthcare settings.
AOI: Trainable multi-agent LLM framework for autonomous cloud diagnosis and SRE automation learning from failed trajectories.
SWE-CI: Benchmark evaluating LLM agent capabilities in continuous integration for long-term codebase maintenance and feature iterations.
Optimizes KV cache in transformers by reducing key dimensionality to log(N) while preserving value information for efficiency.
CRIMSON: Clinically-grounded LLM metric for evaluating chest X-ray report generation on diagnostic correctness and patient safety.
Systematic framework defining boundaries between AI models and AI systems for regulatory and policy compliance purposes.
Amnesia: Adversarial semantic activation steering technique for controlling harmful content generation in large language models.
DUCTILE: Agentic LLM orchestration framework for automating engineering analysis and tool coordination in product development.
ELISA: Interpretable AI agent combining scGPT embeddings with BioBERT for expression-grounded discovery in single-cell genomics.
FRAME: Systematic framework for real-world AI evaluation generating contextual evidence on model behavior in organizational environments.
APEX-Searcher: LLM agent with agentic planning for multi-hop retrieval-augmented generation to enhance complex question answering.
Audio-visual speech enhancement using RL with LLM-based interpretable reward model for perceptual quality optimization.
arXiv: Post-hoc model-agnostic explanation method using perturbation selection for uncertainty-aware surrogate model approximations.
arXiv: Analyzes error sources in global feature effect estimation (PD, ALE plots) for black-box model interpretation.
arXiv: Open-source biomedical knowledge graphs (Pathways, Clinical Trials, Drug-Gene) with AI agent access via Samyama database.
arXiv: Addresses LLM limitations in private-library code generation; shows API documentation retrieval alone is insufficient.
arXiv: HindSight framework evaluates LLM-generated research ideas by matching against future publications and citation impact.
arXiv: Analyzes how wider beam search can hurt LLM output quality due to overestimation bias in noisy scoring.
arXiv: PokeAgent benchmark for multi-agent AI decision-making with partial observability, game theory, and long-horizon planning in Pokemon RPG.
arXiv: Physics-informed neural networks for simulating EUV electromagnetic wave diffraction in lithography. Domain-specific neural networks.
arXiv: Analyzes tokenization design choices for foundation models trained on structured electronic health records.
Reinforcement learning framework extending RLHF with multi-dimensional contextual rubric rewards and alternating optimization.
Inference-time steering mechanism for frozen LLMs using adaptive prompt routing to enable evolving safety alignment without retraining.
Prototype-based OOD detection method with dynamic prototype count adaptation based on category complexity.
Federated learning framework combining knowledge graphs and temporal transformers for early sepsis prediction across multi-center ICUs.
Study of Gini Index role in detecting and debiasing class accuracy disparities in prompt-based classification tasks.
Defense mechanism against steganographic collusion in multi-agent RL using dynamic representational circuit breaking at optimization substrate.
Attribution-guided framework using rank-one model editing to rectify unreliable neural network behavior on non-robust features.
Analysis of transformer training dynamics via spectral edge detection showing parameter updates concentrate in few coherent directions.
Domain adaptation method for remaining useful life prediction with incomplete degradation trajectories using evidential learning.
Hypergraph neural network approach using Ricci flow to address over-smoothing and improve message passing.
Multi-expert framework with uncertainty guidance for imbalanced sequence learning and minority class detection.
Method bridging learned embeddings and interpretable handcrafted features for temporal event sequences in financial systems.
Metacognitive test-time reinforcement learning framework for unified multimodal models enabling knowledge accumulation across similar prompts.
Physics-grounded multimodal LLM agent combining language models with PDE solvers for scientific reasoning without domain-specific fine-tuning.
Zero-shot forecasting method for time series with exogenous variables using prior-fitted networks.