Evian: Towards Explainable Visual Instruction-tuning Data Auditing
Evian: Explainable auditing framework for vision-language model training data, identifying nuanced semantic flaws and logical errors.
Evian: Explainable auditing framework for vision-language model training data, identifying nuanced semantic flaws and logical errors.
Multi-agent LLM system for research idea generation using combinatorial innovation and iterative search to reduce repetition and increase novelty.
Cross-lingual quality classifiers for multilingual pretraining data selection, leveraging embedding space consistency across languages.
LayerTracer: Framework for analyzing hierarchical representations and robustness across diverse LLM architectures (Transformer, Mamba, GateDeltaNet).
Study of emergent social dynamics in LLM agents playing multi-round deception game with memory retention, analyzing reputation formation.
Comparison of discretization strategies for Vision Mamba state space models to improve temporal fidelity in dynamic visual environments.
RSRCC: Remote sensing benchmark with 126k questions for change detection and explanation in satellite imagery using vision-language models.
GRPO-VPS: Improves Group Relative Policy Optimization with verifiable process supervision for better reasoning in LLMs through refined credit assignment.
Analysis of trustworthiness issues in Vision-Language Models, arguing current VLMs fail to faithfully synthesize multimodal data and proposing solutions.
ORPHEAS: Bilingual Greek-English embedding model optimized for retrieval-augmented generation, addressing morphological complexity and domain-specific terminology.
arXiv paper on decision-making frameworks augmented by machine intelligence for high-consequence scenarios.
arXiv paper: StormNet uses graph neural networks and convolution for spatio-temporal storm surge forecasting with bias correction.
arXiv paper: QuanForge mutation testing framework for quantum neural networks addressing stochastic factors and interpretability challenges.
arXiv paper benchmarking omnimodal notation processing for music intelligence across auditory, visual, and symbolic domains.
Data-centric framework using parameter-efficient fine-tuning for adapting LLMs to multiple languages with reduced interference.
Study on prompt optimization and judge selection for LLM-as-a-Judge evaluations in legal question answering tasks.
Training method for adapting smaller LLMs to agentic tasks by generating supplemental text without retraining large models.
Semantic stratification method for improving retrieval evaluation in RAG systems by addressing bias in evaluation set construction.
Investigation of human-like working memory constraints in Transformer architectures for improved learning under data scarcity.
Multidimensional evaluation of general-purpose and medical domain LLMs on clinical communication standards.
Benchmark for evaluating Olympiad-level multi-image reasoning in vision-language models with distributed evidence.
Analysis of how different language model architectures learn similar periodic number representations in Fourier domain.
Open-source framework for identifying vulnerabilities and evaluating security of AI systems across critical domains.
Method for building person-specific generative agents using LLMs grounded in self-report data for general-purpose behavioral simulation.
Survey examining scaling strategies for LLM reasoning including multi-agent collaboration and impact on model performance.
Novel framework using multi-armed bandit optimization to identify influential context segments in retrieval-augmented generation systems.
Open-source framework for building general AI agents with deep research and autonomous capabilities without relying on proprietary APIs.
AstaBench provides rigorous benchmarking suite for evaluating AI agents on scientific research tasks including literature review, experiments, and data analysis.
FELA applies multi-agent evolutionary system to automated feature engineering for industrial event logs, handling scale, dimensionality, and temporal complexity.
Improves test-time scaling of multi-step LLM reasoning by probing internal states to efficiently verify intermediate steps without expensive Process Reward Models.
REST benchmarks evaluate cross-modal inconsistency in MLLMs using samples with identical semantic content in image, text, and mixed modalities.
Proposes epistemic constitution framework for regulating how LLMs form and express beliefs, addressing source attribution bias and implicit epistemic policies.
LEAD addresses no-recovery bottleneck in long-horizon LLM reasoning by analyzing error distribution and showing balanced decomposition prevents irreversible failures on hard steps.
Proposes lightweight external memory system for LLM agents using small language models, reducing overhead while maintaining accuracy for cross-turn consistency in long-horizon tasks.
Uses agentic LLM feedback frameworks for planning domain generation from natural language, improving quality of generated domains through iterative model space reasoning.
Proposes standardized pipeline for log analysis in AI systems with Python code examples, establishing best practices for understanding model capabilities and evaluation integrity.
RAG-KT applies LLM-based retrieval-augmented generation to knowledge tracing for student performance prediction, improving transfer and interpretability across learning platforms.
QuarkMedSearch builds a long-horizon deep search agent for Chinese medical domain, systematically exploring data construction, training strategies, and evaluation for vertical agentic systems.
TREX automates LLM fine-tuning via multi-agent tree-based exploration, orchestrating researcher and executor agents to handle requirement analysis, experimentation, and optimization.
MirrorBench evaluates self-centric intelligence in MLLMs by introducing mirror reflection tasks, addressing gap in embodied AI benchmarks focused on external object understanding.
Applies conformal prediction to VLM-generated guidance in hybrid decision-making, ensuring reliable human-AI collaboration where AI provides guidance rather than decisions.
PersonalHomeBench evaluates foundation models as agentic assistants in personalized smart home environments through iterative benchmark construction with rich household states.
Analyzes modality preference bias in omni-modal LLMs using a conflict-based benchmark, revealing that native unified architectures still exhibit systematic modality preference patterns.
Studies structural limits of enforcement mechanisms in autonomous agent systems, showing how behavioral drift can become invisible to enforcement engines operating at wrong abstraction levels.
CoSearch jointly trains reasoning and document ranking via RL for agentic search, optimizing both agent reasoning and retrieval components together rather than treating retrieval as fixed.
LiteResearcher proposes scalable RL framework for training deep research agents, decoupling real-world search from training.
Explicit Trait Inference improves LLM-based multi-agent coordination by tracking partner warmth and competence characteristics.
Studies fairness and bias in LLMs during role-playing scenarios, assessing whether social biases emerge in simulated roles.
BatchLLM optimizes batched LLM inference throughput using global prefix sharing and throughput-oriented token batching techniques.
Proposes recency bias mechanism for transformer attention in time-series forecasting using attention score reweighting.