Understanding Self-Supervised Learning via Latent Distribution Matching
Provides unifying theoretical framework for self-supervised learning via latent distribution matching principle.
Provides unifying theoretical framework for self-supervised learning via latent distribution matching principle.
Combines physics constraints with rectified flow for reconstructing spatiotemporal PDE-governed fields from sparse measurements.
Presents HeadQ, a KV-cache quantization method optimizing for model-visible distortion in LLM inference.
Studies parameter identifiability in deep ReLU networks using weighted polyhedral complexes framework.
Proposes few-step generative modeling framework using cumulative flow maps for probability space transport, inspired by physical dynamics.
Information plane analysis of binary neural networks using discrete activations to overcome MI estimation challenges in deep network training dynamics.
ELAS combines low-rank training and 2:4 activation sparsity for efficient LLM pre-training with reduced memory and computation.
Uni-OPD framework unifies on-policy distillation theory, identifying bottlenecks in consolidating expert models into single student models.
Fine-tuned LLMs for predicting neural network performance across datasets in AutoML frameworks, bridging code generation and performance reasoning.
GNN-based method for hierarchy-aware knowledge graph embeddings using semantic loss from ontologies, applied to yeast gene deletion phenotype prediction.
Evolutionary Dynamic Loss framework learns transferable classification losses using synthetic data and ranking-consistency objectives without real samples.
Re-examines LoRA rank thresholds for LLM fine-tuning, questioning prior NTK analysis and its practical implications for cross-entropy loss.
Introduces Nora, a matrix-based optimizer for LLMs combining Muon-like preconditioning with scale-invariance and computational efficiency.
Proposes memory-efficient continual learning for CLIP models using distributed contrastive loss to reduce forgetting with small memory buffers.
Analyzes zeroth-order optimization for LLM fine-tuning, showing adaptive ZO methods offer no convergence advantage over ZO-SGD despite memory overhead.
Uses pretrained MLIP representations as acquisition signals for active learning in machine learning interatomic potentials for chemistry.
Thermodynamic-inspired framework for analyzing LLM stability under uncertainty using entropy-based metrics beyond aggregate accuracy.
Kerimov-Alekberli model applying information geometry and non-equilibrium thermodynamics to AI safety and autonomous system alignment.
Analysis of how CLIP embeddings unexpectedly drive memorization in Stable Diffusion text-to-image models, with categorization of token types.
CreativityBench benchmark evaluating LLM creative problem-solving through affordance-based tool repurposing tasks.
VANGUARD system using multimodal LLMs for video anomaly detection with interpretable reasoning and spatial localization of anomalous events.
Evaluation of same-model self-verification as confidence signal for LLMs compared against likelihood-based baselines on reasoning tasks.
Study of Hebbian Fast-Weight modules in vision transformers enabling rapid within-episode adaptation for few-shot character recognition.
EvoJail method for generating diverse automated jailbreak prompts for LLMs using evolutionary algorithms, designed for adaptability to updated safety-finetuned models.
Framework applying evolutionary methods to analyze and explain LLMs by relating model weights to genotypes and outputs to phenotypes.
EFGPP framework for genotype-to-phenotype prediction combining multiple data sources for complex trait prediction, tested on migraine prediction.
SALO method for detecting jailbreaks by tracing dynamic refusal trajectories in LLM representations using causal analysis, robust against adversarial attacks.
MedStruct-S benchmark for semi-structured information extraction from OCR clinical reports with tasks for key discovery, QA, and key-value pair extraction.
Method to detect whether LLM can answer queries before generation by measuring hidden state geometry deviation from answerable reference sets, tested on Llama, Qwen, and Mistral models.
Sparse Memory Finetuning (SMF) method for adapting LLMs while reducing catastrophic forgetting by selectively updating memory rows during training. Compared against LoRA and full finetuning.
RLDX-1 robotic policy model extending Vision-Language-Action models with motion awareness, memory, and physical sensing for complex real-world tasks.
Imbalanced classification framework addressing underrepresented classes under operational capacity constraints for rare detection tasks.
Contrastive regularization method for accent-robust ASR using supervised contrastive learning as auxiliary CTC fine-tuning objective.
Treats coordination as architectural layer for multi-agent LLM systems, addressing 41-87% production failure rates from coordination defects not capability gaps.
Proves sharp dimension-accuracy tradeoffs in embedding representations showing accuracy collapse when embedding dimension mismatches ground truth.
APEX predicts popularity for AI-generated music using multi-task learning with aesthetic quality metrics on new AI music landscape.
GRPO-TTA applies Group Relative Policy Optimization to test-time adaptation of vision-language models via prompt refinement and RL.
Demonstrates LLM safety bypasses using mathematical encoding (set theory, logic, quantum mechanics) achieving 46-56% attack success across eight models.
Proposes capability taxonomy for time series reasoning models applied to financial domain with 2x2 classification framework for entity and temporal analysis.
Formalizes memory poisoning attacks on retrieval-augmented LLM agents using game theory, proposing MEMSAD detection method with unified evaluation framework.
Systematic evaluation of Graph-Tokenizing LLMs questioning effectiveness of treating graph data as prefix tokens for LLMs.
SURE-RAG framework for selective retrieval-augmented generation that verifies evidence sufficiency before answering.
Workspace-Bench 1.0 benchmark evaluates AI agents on file-dependent workspace tasks with real-world dependencies at scale.
Convergent-Divergent Routing steers LLM inference toward desired ethical frameworks by gating transformer pathways at branch points.
Agentic-imodels introduces an autoresearch loop that evolves data-science interpretability tools designed for agent understanding rather than human interpretation.
Manokhin Probability Matrix diagnostic framework separating classifier reliability and discriminatory power using 2x2 grid methodology.
Conformal predictive self-calibration framework for multimodal learning on low-quality data with modality imbalance and noise.
Framework that formulates prompt steering as activation steering for LLMs, investigating whether distillation can close performance gaps.
Task vector arithmetic for combining independently fine-tuned bioacoustic encoders without sharing data across taxa and institutions.
Conditional diffusion sampling approach for sampling from unnormalized multimodal distributions with limited density evaluations.