Visual Prompt Discovery via Semantic Exploration
Visual prompt discovery method to diagnose and mitigate LVLM perception failures through semantic exploration.
Visual prompt discovery method to diagnose and mitigate LVLM perception failures through semantic exploration.
Vision-language process reward models with explicit visual premise verification for reliable step scoring in reasoning.
Genetic programming with surrogate models for dynamic multi-mode project scheduling with simulation-based optimization.
VisBrowse-Bench benchmark for evaluating visual-native search in multimodal browsing agents using MLLMs.
End-to-end framework using Speech LLMs for spoken question answering with attention-guided evidence grounding.
Human-centered architecture for integrating LLM-based cognitive assistants into manufacturing quality management systems.
Security research on sentiment steering attacks targeting RAG-enabled large language models and LLM robustness.
Machine learning pipelines for radio astronomy data processing with explainability focus on automating configuration.
YOLO-based deep learning for automated wasp identification with explainable AI integration for taxonomic classification.
DynamicGate MLP conditional computation framework using learned structural dropout and input-dependent gating for efficiency.
FederatedFactory zero-dependency framework for federated learning in non-IID scenarios using generative one-shot learning.
Physics-guided diffusion framework for full-waveform inversion combining score-based generative models with wave-equation simulations.
Fanar 2.0 Arabic generative AI platform built on 256 H100 GPUs at QCRI with sovereign infrastructure and data pipelines.
Study identifying flaws in LLM benchmarks for Icelandic, highlighting issues with synthetic and machine-translated evaluation data.
PlotTwist creative plot generation framework using small language models with specialized training for narrative coherence.
Method for adding persistent memory to frozen encoder-decoder LLMs via trainable adapters in continuous latent space.
IndexRAG approach for multi-hop question answering that performs cross-document reasoning at indexing time using bridge entities.
SF-Mamba state space model for vision addressing non-causal patch interactions with improved computational efficiency.
SlideFormer system for fine-tuning large language models on single GPU via asynchronous engine and sliding window approach.
EngGPT2-16B open Italian LLM trained on 2.5T tokens, efficient inference with performance comparable to larger models.
Multi-agent reinforcement learning approach for managing delayed channel state information in multi-satellite communication systems.
Unlearning method for one-step generative models using unbalanced optimal transport for safer image generation.
LenghuSky-8 millisecond-resolution network dataset for time series foundation models with high-frequency data.
FEAT foundation model with linear complexity for structured data in healthcare, finance, and e-commerce with improved scalability.
DanceHA multi-agent framework for document-level aspect-based sentiment analysis, extracting ACOSI tuples from documents.
CompDiff uses hierarchical compositional diffusion to generate fair medical images across demographic groups and intersections.
EmoLLM framework integrates appraisal-grounded cognitive-emotional reasoning into LLMs for contextually appropriate responses.
Analysis of human-LLM chat logs characterizing delusional spirals and negative psychological effects from extended chatbot interactions.
Manifold-Matching Autoencoders regularize autoencoders by aligning pairwise distances between latent and input spaces.
Research on classifying malicious AI agent skills using repository context to improve detection in skill marketplaces.
REFORGE reveals vulnerabilities in image generation model unlearning through multi-modal adversarial attacks in black-box settings.
BATQuant proposes outlier-resilient MXFP4 quantization via learnable block-wise optimization for deploying MLLMs and LLMs on accelerators.
Analysis of multimodal LLM-generated natural language explanations for face verification on unconstrained face images.
Omanic, a benchmark for step-wise evaluation of multi-hop reasoning in LLMs with step-level annotations for diagnosing failures.
Investigation of linguistically related language guidance for LLM translation in low-resource settings without large parallel data.
Study of emergent AI agent communities on platforms, analyzing 167k+ agents learning from each other without researcher intervention.
Kestrel, a training-free method for mitigating hallucinations in large vision-language models using grounding and self-refinement.
World action models for embodied control that eliminate test-time future imagination while maintaining action performance.
Resource-aware LLM-based agent reasoning for embodied robots using reinforcement learning to balance computation and action execution.
In-context learning improvement for vision-language models using retrieved counterfactuals for better visual reasoning.
SpecMoE mixture-of-experts foundation model for cross-species EEG decoding with spectral-temporal fusion.
Formal model for selecting statements that find common ground across diverse preferences using generative AI.
TurnWiseEval benchmark and analysis of multi-turn vs single-turn LLM capabilities with step-level evaluation.
InCoder-32B, a 32B code foundation model optimized for industrial programming tasks with hardware semantics and resource constraints.
Cross-embodiment dexterous grasping policy enabling zero-shot transfer across different robot hand morphologies without retraining.
Behavior tree planning for robot manipulation using context-aware grounding to automate controller design without extensive manual effort.
Study examining reasoning mechanisms in diffusion-based video models, challenging chain-of-frames assumptions about how reasoning emerges.
Comprehensive survey of LLM reasoning covering inference scaling, learning to reason, and agentic systems as key advancement areas.
CHARM method calibrates reward models using Chatbot Arena scores to mitigate model preference bias and reward hacking in RLHF.
IMAIA interactive maps assistant enables natural language interaction with vector maps and satellite imagery with geospatial intelligence.