A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
Theoretical analysis of delayed loss spikes in batch-normalized linear models during neural network training.
Theoretical analysis of delayed loss spikes in batch-normalized linear models during neural network training.
Research on privacy-preserving neural network inference using homomorphic encryption for batch processing of encrypted inputs.
Lorentz hyperbolic space semantic segmentation framework with uncertainty quantification and numerical stability improvements.
EFDiff framework uses Prithvi-EO-2.0 foundation model with diffusion for land surface temperature super-resolution.
GRAIL framework enables autonomous grounding of relational concepts in neuro-symbolic RL agents without manual definition.
EasyVideoR1 extends reinforcement learning from verifiable rewards to video understanding in multimodal models.
Freshness-aware prioritized experience replay improves sample efficiency in off-policy RL for LLM/VLM post-training.
Talk2AI longitudinal study measuring LLM persuasiveness on societal topics via psychological and reasoning dimensions.
Bolzano multi-agent LLM system produces novel mathematics results through parallel provers and verifier with persistent knowledge base.
SPS method improves exploration in RL for LLMs by addressing probability squeezing limitations in reasoning model training.
Self-play framework for improving LLM code reasoning via semantic equivalence validation using formal verification and Liquid Haskell proofs.
CAAF framework enforces determinism in LLM-based agentic workflows for safety-critical systems using convergent execution patterns.
Develops trajectory-restricted framework for analyzing linear convergence of first-order optimization methods with geometry-aware bounds.
Proposes stability-weighted decoding for diffusion language models using temporal token stability to improve parallel text generation.
EvoComp reduces visual token count in multimodal LLMs while preserving accuracy using semantic-guided evolutionary compression.
Develops ML-based classifier to automatically identify plasma regions around Mars (solar wind, magnetosheath, magnetosphere) from mission data.
Presents Local Inconsistency Resolution algorithm for learning and inference in probabilistic models using Probabilistic Dependency Graphs.
FlowRefiner uses flow matching for iterative refinement of 3D turbulent flow neural PDE solvers with improved autoregressive prediction.
SynthFix combines LLMs with symbolic AI and compiler feedback for automated code vulnerability repair using neural-symbolic hybrid approach.
Proposes bilinear input modulation for Mamba SSMs to improve memory retention and computational capacity using Koopman bilinear forms.
Security analysis of bit-flip vulnerabilities in shared KV-cache blocks within vLLM's prefix caching for LLM serving systems.
RoTRAG: retrieval-augmented generation method for multi-turn dialogue harm detection using explicit normative principles and rules of thumb.
ARMove: agentic reasoning framework for human mobility prediction using LLMs with improved interpretability, iterative learning, and transferability.
Hybrid ensemble method combining Chain-of-Thought and Program-of-Thought reasoning to achieve self-consistency with only two LLM samples.
SPECTRA: supervision-free framework using cold-start reinforcement learning to improve vision-language model agents' visual perception and tool use.
Semantic search engine for retrieving mathematical knowledge across millions of documents to ground AI mathematics systems.
ONTO: token-efficient columnar notation for serializing operational data to LLMs, reducing JSON overhead by ~80% in IoT sensor datasets.
Diffusion model approach for identifying parameters in nonlinear spatiotemporal systems with improved robustness in turbulent regimes.
Study showing LLM-based agents fail to incorporate unexpected environmental observations into reasoning across three benchmarks, limiting their adaptability.
Research on characterizing language model skills through model-native representations rather than external taxonomies for intervention on model behavior.
Theoretical and practical methods for improving reproducibility in ML estimation by controlling random seed instability via bagging approaches.
GLMTest framework for LLM-based test generation using program structure awareness to target high-risk branches and improve bug discovery.
Validity screen framework for LLM confidence signals evaluated on 20 frontier models, showing predictive validity for selective prediction performance.
Controlled developer study measuring security training effectiveness on LLM-assisted Java Spring Boot implementation with within and between-subject design.
Token-level selective unlearning for LLMs using importance-weighted forgetting to maintain model utility while removing harmful information.
Dataset generation method using adversarial competition between attackers and defenders to create diverse, high-quality conversational training data.
Analysis of AI-assisted code generation failures as control problems, proposing governance layers for tracking structural commitments in human-AI collaboration.
Framework for off-policy learning in constrained MDPs with layered dependency structures for engineering simulation repair tasks.
Zero-shot ship detection in SAR imagery using foundation models and YOLO, demonstrating prompting approaches for specialized domain tasks.
Study examining whether LLM failures in formal linguistic competence stem from architecture limits or data scarcity via controlled pretraining experiments.
Mathematical connection between binarized neural networks and Sugeno integrals, showing BNN inference as rule-based systems.
Research on using LLMs for molecular generation under chemical constraints, framing creativity as a functional requirement for exploring large chemical spaces.
Depth Registers method for W4A4 post-training quantization in language models, reducing error from 1727x to 14x on SwiGLU models.
Dataset distillation technique using soft label pruning and quantization to reduce storage overhead on large-scale ImageNet datasets.
Activation quantization technique (AQPIM) for running large LLMs on Processing-in-Memory architectures with reduced memory footprint.
Distributional off-policy evaluation using deep quantile process regression for estimating return distributions in reinforcement learning.
R package mlr3torch for deep learning on tabular data and tensors, integrating torch with mlr3 ecosystem for classification and regression.
Diffusion model pipeline for zero-shot object grounding in remote sensing imagery, combining diffusion with SAM segmentation models.
Mixture-of-Experts applied to object detection for domain-specialized models with improved interpretability over standard ensembles.
Hebbian learning approach for continual audio classification using kernel plasticity to balance learning new information with retaining previous knowledge.