CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
CR-Net enables parameter-efficient LLM pre-training using cross-layer low-rank structure with reduced memory and computation.
CR-Net enables parameter-efficient LLM pre-training using cross-layer low-rank structure with reduced memory and computation.
Diffusion-based world model for offline reinforcement learning that jointly generates actions, states, and rewards.
LiLAW method that dynamically adjusts training sample weights based on difficulty for noisy data training.
Theoretical analysis of linear models for time series forecasting, examining robustness and interpretability.
Framework enabling LLMs to perform multi-step constrained reasoning by satisfying symbolic constraints in planning tasks.
Research on expert pruning vs merging for compressing Mixture-of-Experts models, showing pruning superior for generative tasks.
Latent-Augmented Discrete Diffusion: learnable auxiliary channels for improved few-step language generation.
Study of architectural factors and scaling laws trade-offs for inference-efficient LLMs.
Loopholing: mechanism preserving distributional information in discrete diffusion for parallel decoding.
Theoretical analysis of sample complexity in differentially private policy optimization for RL.
Sharpness-guided group relative policy optimization for improving LLM reasoning via reinforcement learning.
Radial compensation technique addressing radius distortion in chart-based generative models on Riemannian manifolds.
Open-access ionospheric dataset integrating diverse space weather observations for ML-based forecasting.
Energy scaling laws for diffusion models predicting computational demands across configurations.
FOAM: memory-efficient LLM training via blocked state folding to reduce Adam optimizer overhead.
Training-free asynchronous reasoning for LLM agents enabling real-time responses without sequential thinking bottlenecks.
Benchmark for evaluating attribute discrimination in infant-scale vision-language models on color, size, and texture.
Mechanistic study of feature evolution during training in infinite-depth ResNets under depth-μP scaling.
Deep Delta Learning: transformer residual update rule enabling selective rewriting of content while preserving identity paths.
Flow-based processes for LLM regression tasks addressing error cascades and computational intensity in time-series and text-conditioned predictions.
Theoretical study of transformer forward passes as interacting particle systems with gradient flow interpretation and mean-field limits.
Proposes parameter-level gradient analysis to mitigate catastrophic forgetting during knowledge injection in large language models.
Mechanistic study showing Prior-Data Fitted Networks learn structured spectral representations extractable as explicit kernels for amortized Bayesian inference.
Theoretical analysis of gradient descent on Kolmogorov-Arnold Networks with bounds on optimization, generalization, and differential privacy properties.
Cascaded flow matching approach for generative modeling of heterogeneous tabular data with mixed discrete and continuous features.
Theoretical analysis of cooperative multi-agent reward-free exploration in tabular finite-horizon MDPs with phased learning framework.
Adaptive conformal prediction method for uncertainty quantification under distribution shift in autonomous systems without exchangeability assumptions.
Proposes Preserve-Then-Quantize method for post-training quantization in LLMs, balancing rank budgets between preserving intrinsic structure and error reconstruction.
Theoretical analysis of DeepBern-Nets using learnable Bernstein polynomial activations with guarantees on approximation error and parameter efficiency.
Introduces EntRGi for reward guidance in discrete diffusion language models, enabling test-time adaptation without differentiating through discrete tokens.
Unified framework combining Bayesian optimization and experimental design via active inference for hybrid learning and optimization workflows.
Extends flow matching to offline reinforcement learning with discrete actions and multiple objectives using continuous-time Markov chains.
Proposes PEPO, single-step direct preference optimization algorithm using pessimistic ensemble to avoid over-optimization in LLM preference learning without explicit reward models.
Introduces group causal counterfactual policy optimization to improve LLM reasoning generalization by crediting sound reasoning processes, not just correct answers.
Proposes diffusion-inspired probabilistic transformer architecture for uncertainty calibration in risk-sensitive applications.
Systematic evaluation of chemical language model scaling across model size, dataset size, and compute on molecular property prediction downstream tasks.
Uses message-passing neural networks as heuristics for combinatorial optimization problems, combining learning with approximation algorithm guarantees.
TS-Haystack: benchmark for evaluating time-series language models on retrieval and reasoning over long temporal contexts across 10 tasks.
Research showing test-time training with KV binding is mathematically equivalent to learned linear attention, contradicting memorization interpretation.
Zatom-1: multimodal foundation model unifying generative and predictive learning for 3D molecules and materials across domains.
Research on feature attribution via optimal transport flows, studying path ambiguity in counterfactual explanations.
Research extracting fair bias-agnostic subnetworks from vanilla trained models without additional unbiased data.
Data Agent learns dynamic data selection via end-to-end optimization to prioritize informative samples and accelerate training.
Data-driven integration kernels framework for interpretable nonlocal operator learning in climate and weather modeling applications.
FlashSampling fuses categorical sampling into LM-head matmul for fast exact sampling without materializing logits in large-vocabulary decoding.
Challenge on membership inference attacks against synthetic tabular data generated by diffusion models for privacy evaluation.
Research on distributed reinforcement learning addressing negative learning from high-surprisal data in stale or mismatched actor settings.
Row-momentum normalized preconditioning optimization method for scalable neural network training capturing curvature information efficiently.
Research on uncertainty-guided tree search decoupling exploration from policy optimization for hard exploration RL problems.
Research unifying classifier-free guidance with alignment objectives in diffusion models via inter-class likelihood-ratio maximization.