Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance
Credit assignment via Wasserstein distance on hidden states for fine-grained supervision in reinforcement learning with verifiable rewards.
Credit assignment via Wasserstein distance on hidden states for fine-grained supervision in reinforcement learning with verifiable rewards.
Rabtriever: efficient rationale-based retrieval using on-policy distillation from LLM-based generative rerankers with dual encoding.
Layered security framework for agentic AI systems addressing threats across planning, memory, tool use, and multi-agent coordination components.
Mechanistic analysis of chain-of-thought reasoning in LLMs via activation patching on GSM8K, showing hidden states contain task-relevant information.
Empirical evaluation of locally deployed LLMs (LLaMA 3.2, Mistral) for bug detection in Python code in resource-constrained environments.
Parametric memory head for generative retrieval enabling dynamic document updates. Addresses limitation of static parametric encoding in GenIR systems.
Reproduces and tests Planning Ahead in Generative Retrieval (PAG) method. Addresses beam search pruning vulnerability in document ranking systems.
Federated learning approach combining adaptive quantization and differential privacy for non-IID data. Reduces communication overhead while preserving privacy.
Explainable AI framework for automated knee osteoarthritis grading from X-rays. Decomposes predictions into interpretable radiographic features.
Second-order optimization method for online learning with streaming data. Reduces complexity of inference procedures with accelerated sketching.
Federated learning protocol for cross-institution fraud detection with privacy guarantees. Addresses scalability and integrity in distributed financial systems.
Evaluates clinical harm from safety training in LLM mental health agents. Tests 4 models on therapy scenarios; over 1/3 show psychological deterioration.
Probes Vision Transformer layers for spatial understanding encoding. Shows boundary and depth structure becomes decodable at different depths.
Benchmark suite of Reddit datasets for mental health detection via NLP. Consolidates existing task-specific corpora into reproducible resource.
Empirical study of security risks in multi-agent AI systems. Shows architectural decisions create attack surfaces beyond individual agent robustness.
Detects misaligned reasoning in continuous thought models where LLMs reason in latent space. Addresses safety concerns in expressive reasoning systems.
Uses spatial transcriptomics data as alternative supervision for deep learning nuclei segmentation and classification in pathology images.
Evaluates granularity effects in graph-based AML systems on blockchain. Measures how scoring levels affect compliance queue composition.
Uses GNNs and Bayesian methods for inferring structural health states in civil infrastructure. Domain-specific inverse problem solving.
OCSR system translating molecular images to SMILES strings using closed-loop training with minimum risk. Addresses real-world chemical structure recognition challenges.
Study showing LLMs collapse philosophical diversity compared to human panels. Evaluates 7 proprietary and open-source models on philosophy task alignment.
RouteNLP routes LLM queries across model portfolios to minimize inference costs while maintaining quality. Uses conformal cascading and distillation to handle diverse NLP workloads efficiently.
ComplianceNLP: knowledge-graph-augmented RAG system for detecting regulatory compliance gaps in financial institutions.
TimingLLM: two-stage retrieval-augmented LLM pipeline for pre-synthesis timing prediction from Verilog code.
First grammatical error correction corpus for Romanian with 10k sentence pairs and adapted ERRANT evaluation toolkit.
Hardware-efficient implementations of Softmax and LayerNorm for Transformer models on edge devices with NLP and generative AI applications.
FlowPlace: generative model using flow matching for chip placement optimization with mask-guided generation and real training data.
Practical guide to information-theoretic measures in AI including entropy, cross-entropy, mutual information, and integrated information.
Benchmarks quantum machine learning approaches (QPINN and QRC) for chaotic time-series prediction on NISQ devices.
Multimodal model predicting video-induced pleasure through cognitive appraisal variables from social media content.
FAIR_XAI framework improving fairness and explainability in Vision-Language Models for mental health wellbeing assessment.
TAS-AI: hybrid active learning framework for autonomous spin wave spectroscopy combining model-agnostic and physics-informed approaches.
Investigates heuristic collapse in frontier LLMs when deployed for high-stakes advice in medical, legal, and financial domains.
Deep learning approach for learning turbulence closures in large eddy simulations using differentiable physics and neural networks.
ZenBrain: a 7-layer neuroscience-inspired memory architecture for AI agent systems integrating consolidation and reconsolidation principles.
Comparative study of lightweight deep learning models for mammographic lesion segmentation in resource-constrained environments.
Analysis of representational curvature in LLMs showing how trajectory straightening relates to next-token prediction and behavioral uncertainty.
Reinforcement learning approach for e-commerce product mapping using LLM and multi-agent frameworks to handle noisy listings with promotional keywords.
Adaptive-distribution randomized neural networks framework for PDEs that learns optimal hidden-layer sampling distributions instead of heuristic selection.
Step-level advantage selection approach for efficient LLM reasoning that stabilizes performance while reducing inference-time computation and reasoning trace length.
Performance analysis of multi-agent LLM tutoring systems across different throughput tiers, measuring latency and cost impact of agent specialization.
PEPS method improving implicit neural representations through positional encoding projection sampling for neural fields and texture compression.
Minimum divergence framework for weighting and averaging probabilistic predictions from multiple statistical and machine learning models.
Process Reward Models for agentic data analysis that supervise LLM agents through process-level feedback, addressing silent errors and logical flaws in dynamic tasks.
Extension of non-Euclidean neural quantum states using hyperbolic RNNs including Poincaré and Lorentz variants for quantum systems.
Theoretical analysis of primitive recursion characterizations across recurrent neural networks, polynomial ODEs, and discrete maps using bounded ReLU iteration.
arXiv technical report on Summary Attention mechanism to reduce quadratic complexity of softmax attention in long-context LLMs, enabling better performance for code agentic intelligence and semantic reasoning.
Survey on split learning techniques for privacy-preserving LLM fine-tuning across distributed client-server architectures.
arXiv paper on hardware-aware neural architecture search with low-precision constraints for spaceborne edge AI deployment.
arXiv paper presenting MIMIC, a generative multimodal foundation model trained on aligned biomolecule data across sequence, structure, and regulatory modalities.