Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions
Study of hidden bias in instruction-tuned LLMs showing fair outputs mask biased internal representations affecting high-stakes decisions.
Study of hidden bias in instruction-tuned LLMs showing fair outputs mask biased internal representations affecting high-stakes decisions.
Unified data mixing method for language models across pretraining, continual learning, and adaptation phases.
Authorization framework for autonomous AI agents using proof-derived policies to prevent unsafe actions in sovereign systems.
Kubeflow MLOps integration for adversarial robustness in AI models deployed on Kubernetes.
Transformer-based unified particle simulator for diverse physical phenomena without solver-specific redesign.
SMCEvolve applies Sequential Monte Carlo sampling to LLM-driven program search for automated scientific discovery.
Minerva-Ego benchmark for evaluating egocentric video reasoning with intermediate reasoning step annotations.
Belief Engine for auditable stance dynamics in multi-agent LLM deliberation with inspectable evidence-based belief updates.
Comprehensive analysis of neural activation patterns across six LLM architectures on cognitive tasks.
Study of hidden-state trajectories in reasoning-trained LLMs showing they follow different internal paths during chain-of-thought.
Analysis of when sparse Mixture-of-Experts routing benefits vision models, identifying compute-leverage patterns.
X-SYNTH system for AI agents to synthesize enterprise context from human attention patterns rather than retrieval.
Theoretical analysis proving RoPE positional embeddings lose effectiveness in long-context LLMs.
Distributional process reward model predicting step-level success probability and reliability for reasoning tasks.
Non-autoregressive text generation model using draft-conditioned latent refinement with flow networks.
Multi-agent orchestration framework dynamically switching between parallel and sequential LLM agent collaboration.
Calibration method for LLMs using semantic-level rewards to improve uncertainty estimation in high-stakes tasks.
Study evaluating LLM code generation transfer to unseen programming languages through fine-tuning analysis.
Prompting technique for vision-language models addressing base-new class trade-off in transfer learning.
Offline policy learning framework for contextual bandits optimizing under general risk criteria.
Few-shot LLM approach for triaging online patient inquiries into clinical action categories with minimal labeled data.
Statistical framework improving Concept Activation Vectors for interpretability in deep learning models.
Adaptive autoscaler for container orchestration that learns cold-start duration using EWMA estimation.
Preference optimization framework for flow models addressing intra-group variance decay in RL alignment of generative models.
ML framework for HPC performance prediction using merged execution traces to overcome hardware counter limitations.
Multi-layer cloud intrusion detection system combining LLMs with Q-learning for improved performance in real deployments.
Tool for detecting and interpreting domain shifts in high-dimensional data by identifying anomalies in feature subspaces.
Causal analysis of format inconsistencies in LLM-as-judge scoring using PEAP to investigate internal mechanisms.
RecMem: memory consolidation system for long-running LLM agents that reduces token consumption through lazy consolidation.
Uses property-guided synthesis to reduce LLM inference costs in program synthesis for planning, guiding generation with formal properties.
Asteria: runtime system for scalable LLM training using second-order optimization, decoupling preconditioner state from GPU path.
Combines formal methods with ML for auditing and monitoring AI systems across development lifecycle, enabling compliance checks.
Controlled study of compound LLM agent design trade-offs in adversarial environments, analyzing context, reasoning, and task decomposition.
Technique to characterize language model organization by lesioning parameters, inspired by neuroscience aphasia studies.
arXiv: Framework combining generative AI, smart metering, quantum optimization for energy infrastructure and billing.
arXiv: FORGE—population-based protocol for self-evolving LLM agent memory without gradient updates. ReAct agents with Reflexion.
arXiv: Studies how LLM-mediated communication influences collective opinion formation on social platforms.
arXiv: Data reconstruction attacks against federated learning systems. Privacy and security analysis of FL protocols.
arXiv: Convergence analysis of federated Q-learning in heterogeneous multi-agent environments. Theory for distributed RL.
arXiv: Explores Kolmogorov Superposition Theorem alternatives for neural network design beyond Kolmogorov-Arnold Networks.
arXiv: Tube Loss function for prediction interval estimation in regression. Novel loss function with theoretical guarantees.
arXiv: Semi-supervised learning approach for sparse reward shaping in reinforcement learning. Uses SSL and data augmentation.
GPU scheduling algorithm for LLM inference that manages KV cache memory constraints and minimizes latency under $700k daily inference costs.
Comparison of LLM routing strategies showing simple k-NN outperforms complex learned routers for selecting specialized models.
Comprehensive benchmark for learning surrogate models of stochastic PDEs with complex spatio-temporal dynamics.
Active learning approach for LLM alignment that selectively samples preference annotations, reducing cost while maintaining alignment quality.
Proposes gradient-free neural network training via projection operators and feasibility-seeking, alternative to conventional loss minimization.
Extends double Q-learning to deep RL, improving target bootstrap decoupling in value function estimation.
Combines supervised and reinforcement fine-tuning for LLMs using prefix sampling, addressing trade-offs between behavior cloning and performance gains.
Trains LMs via RL to reason about uncertainty rather than just correctness, improving performance on question answering by penalizing low-confidence outputs.