Simplex-to-Euclidean Bijections for Categorical Flow Matching
Method for learning probability distributions on simplexes via smooth bijections and Aitchison geometry.
Method for learning probability distributions on simplexes via smooth bijections and Aitchison geometry.
UniQL framework combining quantization and low-rank compression with adaptive on-device pruning for edge LLM deployment.
Post-training method achieving 99.6% attention sparsity without performance loss for mechanistic interpretability research.
WebGym: largest open-source environment with 300k tasks for training visual web agents on realistic websites.
Confidence-Variance theory framework for improved pseudo-label selection in semi-supervised learning beyond fixed thresholds.
Method for optimizing interaction between feature alignment and target fitting in cross-modal model fine-tuning.
Framework for learning Hamiltonian flow maps to enable stable large-timestep molecular dynamics simulations.
Data-free early stopping framework for federated learning using task vector growth rate monitoring.
EPIAGENT agentic framework that automatically synthesizes and calibrates epidemiological simulators via iterative program synthesis.
Framework for score-based density ratio estimation addressing path-variance issues in practical training objectives.
Theoretical analysis of phase transitions in neural network feature learning on multi-index models.
Technique to detect misbehaviors and hallucinations in large vision-language models using evidential uncertainty quantification.
Method for training ensemble models that quantify epistemic uncertainty via distributionally robust optimization.
Study of LLM scaling paradox showing larger compressor models can reduce context reconstruction faithfulness despite lower training loss.
Novel sequence architecture using Conformal Geometric Algebra instead of linear operations for improved generalization and interpretability.
On-policy distillation method that aligns student models with teacher logit distributions, theoretically framed as KL-constrained RL with reward extrapolation.
Analysis of multilingual data curation across 13 languages for 20-trillion-token dataset, addressing multilinguality challenges in foundation models.
Introduces Soft Sequence Policy Optimization, advancing LLM alignment methods beyond GRPO with sequence-level importance sampling.
Theoretical analysis connecting random network distillation, deep ensembles, and Bayesian inference for uncertainty quantification in deep learning.
Shows Pass@k optimization can degrade Pass@1 due to prompt interference, revealing trade-offs in LLM fine-tuning for code generation.
AngelSlim comprehensive toolkit for large model compression combining quantization, pruning, distillation, and speculative decoding.
Muon+ optimizer improves LLM pre-training by adding normalization step after gradient orthogonalization.
Knowledge Fusion via SkillPacks enables efficient cross-capability transfer between LLMs for multi-task integration and model compression.
Theoretical analysis of nonlinear attention mechanisms compared to linear regression for understanding interpolation error.
Adaptive hybrid caching optimization for efficient inference in video diffusion transformer models to reduce computational cost.
ConflictScope pipeline automatically evaluates how LLMs prioritize different values when facing conflicting objectives.
Supervised Reinforcement Learning method for training small open-source LLMs on multi-step reasoning tasks, combining SFT and RLVR approaches.
Method for efficient LLM inference via speculative decoding with dynamic tree construction accounting for system variables.
Temporal Sparse Autoencoders using sequential language structure for discovering interpretable features in LLM representations.
Analysis of energy efficiency of small LLMs on local hardware accelerators versus cloud inference for practical deployment.
VLM-Pruner: Token pruning method for vision-language models addressing spatial sparsity and inter-token redundancy.
One-step diffusion samplers using self-distillation and deterministic flow for efficient sampling from unnormalized distributions.
Intelligent agent system for automatically reproducing deep learning bugs by leveraging nondeterminism and environment coupling.
LeanCat: Benchmark of 100 formalized category theory tasks in Lean evaluating LLMs on library-grounded abstraction and theorem proving.
Framework connecting GFlowNets to Markov chain reversibility to control exploration-exploitation trade-off during training.
Study of multi-query synthesis for dense retriever training showing quality-diversity trade-off benefits out-of-domain and multi-hop retrieval.
Evaluation showing GPT-4o lacks causal models of mental states required for true Theory of Mind despite benchmark performance.
Agent with internet access performs at-scale deanonymization of Hacker News and interview participants using LLMs.
LLM4Cov: Offline agent-learning framework using execution feedback from hardware simulators for test generation with high coverage.
Theoretical analysis of how low-precision quantization affects model and data capacities in high-dimensional linear regression.
Method to decompose epistemic uncertainty in Bayesian deep learning by per-class contributions for asymmetric cost classification tasks.
Research on LLM capacity to be persuaded and detect manipulation, testing vigilance and persuasion in high-stakes decision-making contexts.
Research paper on sparse weight editing for multilingual LLM safety alignment across low-resource languages without expensive retraining.
Agentplace tool for building and automating AI agents at scale. Addresses challenges in agent development and scaffolding.
Microsoft Copilot Tasks uses AI agent to autonomously complete user tasks like converting emails to slideshows.
Cloud IDE with agent-driven features, mobile-desktop handoff, built-in shell. Open-source, $5/month after free trial.
Billing platform for MCP tool servers enabling developers to monetize AI agent tools. Stripe integration, TypeScript SDK, live with 6 tools.
Proposal for llm:// URI scheme to standardize LLM connection URLs. Led to draft IETF RFC.
Model Context Protocol server integrating Sharesight portfolio platform with Claude AI assistants.
Claude Code extension providing autonomous C-suite executive decision-making across multiple executive functions.