Distribution-Free Pretraining of Classification Losses via Evolutionary Dynamics
Evolutionary Dynamic Loss framework learns transferable classification losses using synthetic data and ranking-consistency objectives without real samples.
Evolutionary Dynamic Loss framework learns transferable classification losses using synthetic data and ranking-consistency objectives without real samples.
Re-examines LoRA rank thresholds for LLM fine-tuning, questioning prior NTK analysis and its practical implications for cross-entropy loss.
Introduces Nora, a matrix-based optimizer for LLMs combining Muon-like preconditioning with scale-invariance and computational efficiency.
Proposes memory-efficient continual learning for CLIP models using distributed contrastive loss to reduce forgetting with small memory buffers.
Analyzes zeroth-order optimization for LLM fine-tuning, showing adaptive ZO methods offer no convergence advantage over ZO-SGD despite memory overhead.
Uses pretrained MLIP representations as acquisition signals for active learning in machine learning interatomic potentials for chemistry.
Thermodynamic-inspired framework for analyzing LLM stability under uncertainty using entropy-based metrics beyond aggregate accuracy.
Kerimov-Alekberli model applying information geometry and non-equilibrium thermodynamics to AI safety and autonomous system alignment.
Analysis of how CLIP embeddings unexpectedly drive memorization in Stable Diffusion text-to-image models, with categorization of token types.
CreativityBench benchmark evaluating LLM creative problem-solving through affordance-based tool repurposing tasks.
VANGUARD system using multimodal LLMs for video anomaly detection with interpretable reasoning and spatial localization of anomalous events.
Evaluation of same-model self-verification as confidence signal for LLMs compared against likelihood-based baselines on reasoning tasks.
Study of Hebbian Fast-Weight modules in vision transformers enabling rapid within-episode adaptation for few-shot character recognition.
EvoJail method for generating diverse automated jailbreak prompts for LLMs using evolutionary algorithms, designed for adaptability to updated safety-finetuned models.
Framework applying evolutionary methods to analyze and explain LLMs by relating model weights to genotypes and outputs to phenotypes.
EFGPP framework for genotype-to-phenotype prediction combining multiple data sources for complex trait prediction, tested on migraine prediction.
SALO method for detecting jailbreaks by tracing dynamic refusal trajectories in LLM representations using causal analysis, robust against adversarial attacks.
MedStruct-S benchmark for semi-structured information extraction from OCR clinical reports with tasks for key discovery, QA, and key-value pair extraction.
Method to detect whether LLM can answer queries before generation by measuring hidden state geometry deviation from answerable reference sets, tested on Llama, Qwen, and Mistral models.
Sparse Memory Finetuning (SMF) method for adapting LLMs while reducing catastrophic forgetting by selectively updating memory rows during training. Compared against LoRA and full finetuning.
RLDX-1 robotic policy model extending Vision-Language-Action models with motion awareness, memory, and physical sensing for complex real-world tasks.
Imbalanced classification framework addressing underrepresented classes under operational capacity constraints for rare detection tasks.
Contrastive regularization method for accent-robust ASR using supervised contrastive learning as auxiliary CTC fine-tuning objective.
Treats coordination as architectural layer for multi-agent LLM systems, addressing 41-87% production failure rates from coordination defects not capability gaps.
Proves sharp dimension-accuracy tradeoffs in embedding representations showing accuracy collapse when embedding dimension mismatches ground truth.
APEX predicts popularity for AI-generated music using multi-task learning with aesthetic quality metrics on new AI music landscape.
GRPO-TTA applies Group Relative Policy Optimization to test-time adaptation of vision-language models via prompt refinement and RL.
Demonstrates LLM safety bypasses using mathematical encoding (set theory, logic, quantum mechanics) achieving 46-56% attack success across eight models.
Proposes capability taxonomy for time series reasoning models applied to financial domain with 2x2 classification framework for entity and temporal analysis.
Formalizes memory poisoning attacks on retrieval-augmented LLM agents using game theory, proposing MEMSAD detection method with unified evaluation framework.
Systematic evaluation of Graph-Tokenizing LLMs questioning effectiveness of treating graph data as prefix tokens for LLMs.
SURE-RAG framework for selective retrieval-augmented generation that verifies evidence sufficiency before answering.
Workspace-Bench 1.0 benchmark evaluates AI agents on file-dependent workspace tasks with real-world dependencies at scale.
Convergent-Divergent Routing steers LLM inference toward desired ethical frameworks by gating transformer pathways at branch points.
Agentic-imodels introduces an autoresearch loop that evolves data-science interpretability tools designed for agent understanding rather than human interpretation.
Manokhin Probability Matrix diagnostic framework separating classifier reliability and discriminatory power using 2x2 grid methodology.
Conformal predictive self-calibration framework for multimodal learning on low-quality data with modality imbalance and noise.
Framework that formulates prompt steering as activation steering for LLMs, investigating whether distillation can close performance gaps.
Task vector arithmetic for combining independently fine-tuned bioacoustic encoders without sharing data across taxa and institutions.
Conditional diffusion sampling approach for sampling from unnormalized multimodal distributions with limited density evaluations.
Framework measuring how safety and accuracy scale differently in clinical LLMs, showing size increases don't guarantee safer behavior.
Unified framework for tabular generative modeling with loss functions, benchmarks, and multi-objective Bayesian optimization approaches.
Applies symmetric data augmentation with deep deterministic policy gradient for aircraft control. Uses reinforcement learning with sample efficiency.
Method for privacy-preserving LLM inference over encrypted data by approximating Softmax and layer normalization as polynomials for homomorphic encryption.
Studies bias inheritance when LLMs generate synthetic training data, showing how biases propagate and amplify in downstream tasks.
Research on zeroth-order optimization with generative priors using coarse learnability for sample-efficient model-based optimization.
Convergence rate analysis for deep neural network classifiers under hard margin condition.
Graph-conditional diffusion models for joint generation of multi-table relational databases.
Optimal matrix sign computation methods for neural network training via the Muon optimizer.
Selective Jacobi decoding accelerates discrete autoregressive normalizing flows for generative modeling.