The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
Evaluation of sycophancy in LLM-based financial agents, measuring tendency to agree with users over correctness in agentic tasks.
Evaluation of sycophancy in LLM-based financial agents, measuring tendency to agree with users over correctness in agentic tasks.
Open-source framework for building modular, observable, and evolvable multi-agent systems with pluggable components and evolution engine.
Iterative method for dynamically identifying and uncovering factual errors in large language models without extensive human labeling.
Survey of data-centric foundation models in computational healthcare, emphasizing data quality and characterization.
Framework for LLM agents to request clarification when faced with ambiguous or incomplete instructions before executing tools.
Study on optimal language mixture ratios for continual pre-training of Llama-3 70B across multiple languages and domains.
Multilingual study evaluating human ability to detect LLM-generated text across 16 datasets covering 9 languages and domains.
Branch-Merge distillation method for compressing large language models while maintaining performance through knowledge transfer.
Analysis of security challenges in multi-agent AI systems, including collusion, swarm attacks, and privacy breaches in networked agent deployments.
Survey of safety and security threats in computer-using agents, covering vulnerabilities in LLM-based systems that interact with GUIs and applications.
Systematic survey of data balancing methods including oversampling and augmentation techniques for addressing class imbalance in machine learning.
arXiv paper presenting LUNGUAGE benchmark for structured chest X-ray report generation with longitudinal evaluation support.
arXiv paper introducing SpookyBench showing video-language models struggle with temporal-only patterns when spatial information is absent.
arXiv paper on MINOS, a multimodal evaluation model for bidirectional image-text generation assessment.
arXiv paper investigating LLM potential for medical decision-making, comparing treatment problem approaches with evidence-based medicine.
arXiv paper on addressing popularity bias in graph neural network recommender systems using regularized fairness approach.
arXiv paper introducing MedCheck, a lifecycle-oriented benchmark framework for evaluating LLMs in healthcare with clinical fidelity.
arXiv paper proposing Neural Bridge Processes improving expressivity of neural diffusion models with input-conditioned noise.
arXiv paper on neural vertex features for efficient learnable 3D scene representation in neural rendering.
arXiv paper on generating visual navigation instructions from egocentric observations using multimodal reasoning.
arXiv paper on federated learning robustness against Byzantine adversarial attacks with loss-based client clustering.
Analysis of confidence calibration in code-completion LLMs. Evaluates relationship between model confidence and prediction accuracy for code generation.
Empirical study (N=150) examining how LLM conversational agent personality and alignment affect user perceptions in goal-oriented tasks like travel planning.
Hybrid Diffusion framework combining symbolic and continuous planning for long-horizon robotic tasks using diffusion models.
PATCH: Learnable sparsity method for LLM deployment combining unstructured and semi-structured pruning to reduce memory/compute costs while maintaining GPU acceleration.
Auto-ARGUE: LLM-based evaluation tool for citation-backed report generation from RAG systems. Open-source implementation of ARGUE framework.
Information-theoretic framework to detect higher-order structure and emergent coordination in multi-agent LLM systems using data-driven analysis.
Survey of Process Reward Models that evaluate LLM reasoning step-by-step rather than just final outputs. Covers data generation, model building, and test-time applications.
Analyzes entropy collapse in reinforcement learning with verifiable rewards for LLM reasoning via entropy change perspective.
FedPF algorithm for federated learning balancing fairness and privacy using differential privacy and zero-sum game formulation.
EvoDev iterative framework for end-to-end software development using LLM agents inspired by feature-driven development practices.
Systematic evaluation of reference-free factuality metrics for long-document abstractive summarization using LLMs.
Inferix block-diffusion inference engine for world models generating long, realistic videos for agentic and embodied AI.
DIQ-H benchmark and value-guided refinement method for evaluating VLM robustness in adversarial conditions and continuous deployment.
PRAXIS orchestrates LLM-driven agentic workflow for root-cause analysis in cloud incidents using program analysis.
CAPT enables training-free domain adaptation of new-generation LLMs using legacy clinical models via model ensembling.
Exposes selective safety trap showing LLM safety evaluations mask vulnerabilities in underrepresented communities.
AdaFRUGAL automates memory-efficient LLM training by dynamically controlling gradient splitting hyperparameters during optimization.
Glance-or-Gaze uses RL to enable large multimodal models to adaptively focus on relevant image regions for search-augmented queries.
HER framework enables LLMs to simulate personas with human-like reasoning and reinforcement learning for role-playing tasks.
ELIQ label-free framework for assessing quality of AI-generated images without requiring manual labels.
AFlow framework for emotional support conversations using LLMs with intermediate strategy supervision via affective flow modeling.
ReLoop framework addresses LLM-to-code translation failures in optimization problems through structured generation and behavioral verification.
End-to-end policy using continuous flow fields for language-conditioned robot navigation in complex scenes.
TildeOpen LLM: 30B-parameter open-weight model trained on 34 European languages using curriculum learning for equitable linguistic representation.
Addresses feature collisions in class-incremental learning using causal analysis to identify spurious correlations and expand features.
Proposes non-Euclidean distance layers to improve harmonic loss function beyond cross-entropy for deep neural network training.
SciMDR introduces synthesize-and-reground framework for constructing scientific multimodal document reasoning datasets balancing scale, faithfulness, and realism.
Adaptive Layerwise Perturbation unifies off-policy corrections for LLM reinforcement learning to address policy staleness and training-inference mismatch.
Woosh is Sony AI's open sound effects foundation model with audio encoder/decoder and text-to-audio generation capabilities.