High-Rate Quantized Matrix Multiplication II
Quantized matrix multiplication technique for weight-only post-training quantization of LLMs using available covariance information.
Quantized matrix multiplication technique for weight-only post-training quantization of LLMs using available covariance information.
Managed infrastructure system for LoRA post-training and serving of millions of LLM variants using shared base model deployments.
Stateful transformer architecture for efficient streaming inference with persistent KV cache reducing prefill cost to O(|q|) complexity.
Multi-level annotator modeling framework to improve reproducibility of LLM evaluation by accounting for human rater biases and subjectivity.
Theoretical analysis of vector quantization using randomized Hadamard transform for similarity search, federated learning, and KV cache compression.
Entropy-guided approach for efficient test-time scaling of agentic systems on software engineering tasks like code generation and bug fixing.
Review of foundation models for Earth science applications integrating multimodal data for perception and scientific discovery tasks.
Lightweight CNN for brain tumor classification in MRI images. Medical imaging application.
Swin Transformer with federated learning for privacy-preserving UAV image transmission in low-altitude networks.
Improved diffusion posterior sampling method for image restoration using lagged temporal corrections.
Theoretical analysis of computational complexity in diffusion models using statistical field theory.
DocAtlas: Multilingual OCR and document understanding benchmark covering 82 languages and 9 evaluation tasks.
Online conformal prediction method enforcing monotonicity across confidence levels. Uncertainty quantification framework.
Theoretical analysis of inverse temperature scaling in self-attention for long-context stability. ArXiv paper on attention mechanisms.
FePySR: Neural feature extraction framework for symbolic regression. Reduces search space by decomposing expressions into reusable components.
CHAL proposes hierarchical multi-agent debate framework for LLM reasoning, addressing limitations of majority voting and confidence escalation.
Five-layer MLOps architecture for continuous assurance of safety and performance in automated driving systems.
Evaluates whether LLM student simulators faithfully replicate student misconceptions during interaction with AI tutors.
ISOMORPH is a public digital twin and benchmark for supply-chain forecasting with configurable parameters and modular topology.
Develops calibration diagnostics for pseudo-labelled regression using confidence thresholding on classifier outputs.
REALISTA formulates hallucination elicitation in LLMs as constrained optimization finding semantically coherent adversarial prompts.
GraphIP-Bench evaluates model-extraction attacks on graph neural networks and tests ownership defenses against surrogate model theft.
Studies persona-model collapse in emergent misalignment of LLMs fine-tuned on narrow harmful data, affecting unrelated prompts.
ChipMATE uses multi-agent reinforcement learning for RTL code generation with self-trained models compatible with vendor security requirements.
Steer-to-Detect method uses internal LLM representations to detect machine-generated text by steering hidden representations for better discriminative power.
Theoretical analysis of weak-to-strong generalization where strong models are fine-tuned on weaker model outputs, studying feature learning under multi-step SGD.
SHM-Agents integrates large language models with specialized algorithms for structural health monitoring, combining generalist-specialist reasoning for engineering applications.
Adaptive conformal prediction method for medical image classification with coverage guarantees on difficult inputs.
Statistical framework for deciding when LLM-enabled generate-verify workflows should release outputs using always-valid inference.
Generative model combining coreset-induced source distribution with hierarchical rectified flow for conditional velocity matching.
Protocol-driven development framework for governing generated software using invariants and evidence beyond natural language specs.
Large-scale benchmark for autoregressive neural population forecasting with structured evaluation metrics beyond correlation.
Proposes watermarking as monitoring primitive for generative models with per-entity attribution keys and detector access.
Supply-chain backdoor attack on diffusion models via PRNG hijacking, with quantum random number defense mechanism.
Benchmark and empirical study using code language models to detect vulnerability-fixing commits across 180k+ commits and 20+ datasets.
Information-theoretic analysis of knowledge distillation generalization using coupled stochastic processes and KL divergence.
Targeted synthetic data generation method using acquisition functions to improve model training data quality.
Adaptive asymmetric adapter (A3B2) for vision-language few-shot learning addressing branch bias between image and text encoders in CLIP-based models.
LoREnc: Training-free encryption framework securing foundation models and LoRA adapters via spectral truncation against IP leakage and model recovery attacks.
Framework for evaluating LLMs as implicit imputers, showing uncertainty should scale with missing information using multiple imputation criteria.
Cryptographically undetectable backdoor attack mechanism for modern neural networks with backdoors hidden in latent space.
Generalized Advantage Grouped Policy Optimization for reinforcement learning on LLM agents, improving credit assignment in multi-turn sparse reward settings.
Agentic AI framework combining LLMs with chain-of-thought reasoning for UAV logistics scheduling and mobile edge computing task allocation.
Information-theoretic approach to visual evidence selection in multimodal RAG systems, reformulating selection by information gain rather than semantic similarity.
Differentiable learning approach for discovering lifted action schemas in classical planning for deterministic MDPs with structural generalization.
Framework for enhancing LLM extrapolation performance through learnable continuous perturbations in embedding space rather than discrete fixed designs.
Parallel multi-turn medical dialogue dataset spanning English and 9 Indic languages with LLM-generated synthetic conversations for healthcare accessibility.
AI module for open-source Wazuh SIEM using MITRE ATT&CK enrichment for multi-step web attack detection.
Test-time self-training method enabling LLM parameter updates during inference to correct misconceptions and adapt to queries.
Multi-stage framework using foundation models for detecting reclaimed slurs in multilingual social media.