Oversmoothing, Oversquashing, Heterophily, Long-Range, and more: Demystifying Common Beliefs in Graph Machine Learning
Survey demystifying graph neural network concepts: oversmoothing, oversquashing, heterophily, and long-range dependencies.
Survey demystifying graph neural network concepts: oversmoothing, oversquashing, heterophily, and long-range dependencies.
Research on KL-regularized policy gradient algorithms for LLM reasoning, comparing forward/reverse KL regularization designs.
Systematic literature review of explanation user interfaces for black-box AI systems and XAI techniques.
FinTagging benchmark for evaluating LLMs on financial information extraction and hierarchical GAAP concept classification.
Automated web app testing system using LLMs and screen transition graphs for test case generation.
Study on persona-driven prompting of LLMs to simulate voting behavior in European Parliament, analyzing progressive bias mitigation.
Strict Subgoal Execution method improves long-horizon planning in hierarchical reinforcement learning through reliable subgoal feasibility.
MCIF benchmark for evaluating multimodal LLMs on crosslingual instruction-following with long-form inputs from scientific talks.
CareerPooler generative AI system for career exploration using pool-table metaphor simulation instead of linear chat interface.
CoSpaDi training-free compression method for LLMs using sparse dictionary learning instead of rigid low-rank approximations.
First watermarking scheme designed for diffusion language models that generate tokens in arbitrary order rather than sequentially.
Prompt optimization framework extended to multimodal LLMs, optimizing visual and textual prompts jointly for improved performance.
pi-Flow modifies flow-based generative models to predict network-free policies for efficient few-step image generation.
VERA-MH automated evaluation framework for assessing safety of AI chatbots in mental health contexts using LLM-based agents.
LRT-Diffusion applies risk-aware sampling to diffusion policies for offline reinforcement learning with statistical hypothesis testing.
VeriStruct framework uses LLMs to automate verification of data structure modules in the Verus verification language.
Semi-Supervised Preference Optimization reduces labeled feedback requirements for aligning language models with human preferences.
PREPO framework improves data efficiency of reinforcement learning for LLMs by leveraging intrinsic data properties during training.
Evaluation of fine-tuned BERT vs LLM prompting for text classification on South Slavic languages, a less-resourced language group.
Wireless foundation models extended to process multiple modalities for improved task performance across varying operating conditions.
Empathetic Cascading Networks multi-stage prompting framework reduces social biases in LLMs through perspective adoption and emotional resonance stages.
Reveals lexical and positional biases in post-hoc feature attribution methods like Integrated Gradients, affecting explanation quality for language models.
Block-Recurrent Hypothesis characterizes Vision Transformer depth as block-recurrent structure, providing mechanistic understanding of ViT computations.
Framework integrating Theory of Mind into robots for inferring human mental states to enhance explainability and predictability in human-robot interaction.
Mixed-methods audit examining alignment between student preferences and AI system capabilities for collaborative academic tasks in CS education.
Analysis of DARPA's AIxCC competition for autonomous cyber reasoning systems leveraging LLMs to discover vulnerabilities in open-source software.
Framework for jointly optimizing data mixture and model architecture configurations during LLM training through co-optimization rather than sequential approaches.
Automated black-box pipeline detects unverbalized biases in LLM reasoning where models hide internal biases in plausible-sounding chain-of-thought explanations.
LoRA-Squeeze compresses LoRA modules through post-tuning and in-tuning methods to simplify rank selection and improve deployment efficiency for fine-tuning LLMs.
SCOPE framework for calibrated pairwise LLM judging with statistical guarantees, reducing miscalibration and systematic biases in evaluations.
VisPhyWorld evaluation framework tests whether MLLMs reason about physical dynamics through code-driven video reconstruction tasks.
Diary study examining how multimodal LLMs assist blind and low vision users accessing visual information through conversational interfaces.
Optimization method for dataset distillation using exploration-exploitation to compress large datasets while retaining model performance.
Proves logit distance bounds representational similarity for discriminative models including autoregressive language models.
Resp-Agent system uses active adversarial curriculum learning for multimodal respiratory sound generation and disease diagnosis.
Theoretical framework using convex conjugate duality to characterize trainability and generalization properties of deep neural networks.
Evaluates LLMs as zero-shot annotators for Bangla hate speech detection, examining reliability and bias in low-resource language settings.
RoboGene uses agentic framework to automatically generate diverse robotic manipulation tasks, addressing data scarcity in VLA pre-training.
Framework for LLM agents to reason about cost-uncertainty tradeoffs when deciding whether to explore environments before committing to answers.
PCAS system enforces deterministic authorization policies in LLM agents for customer service, workflows, and compliance without relying on prompts.
Few-shot LLM classification framework predicts electricity market price spikes using natural language prompts with system state features.
Studies stability of transformer attention-head circuits across model instances to determine if interpretability findings are universal or idiosyncratic.
DeepVision-103K dataset of 103K visually diverse mathematical problems for multimodal LLM reinforcement learning with verifiable rewards.
PETS framework for principled trajectory allocation in test-time self-consistency scaling with sample efficiency optimization.
Geometric analysis of grokking in transformers showing low-dimensional optimization dynamics with PCA of attention trajectories.
LiveClin live benchmark for clinical LLM evaluation using contemporary peer-reviewed cases updated biannually to prevent contamination.
Analysis of how distribution shifts in language models relate to omitted variable bias and mitigation approaches.
Method improving LLM causal reasoning on counterfactual questions via double counterfactual consistency learning.
Efficient inference pipeline achieving IMO-level math reasoning with off-the-shelf models at reduced computational cost.
Fine-tuning diffusion and flow models for tail-aware generative optimization with control over reward distribution.