Addresses reasoning alignment in multi-modal LLMs under concept drift by formulating the problem as constraint satisfaction to handle systematic biases in non-stationary environments.
TokenChain discrete speech chain framework coupling semantic-token ASR with two-stage TTS for joint improvement of speech recognition and synthesis.
DELTA layer-aware dynamic token attention pruning for long-context reasoning models reducing decode latency through selective KV cache sparsification.
STAR dynamic rescheduling algorithm for LLM decode-phase workload balancing addressing SLO violations and OOM failures from variable output lengths.
Cross-paradigm graph backdoor attacks using promptable subgraph triggers that transfer across supervised, contrastive, and prompt-based GNN paradigms.
GD-FPS parameter-efficient fine-tuning method using feedforward parameter selection to avoid gradient computation overhead while adapting pre-trained models.
P3-LLM heterogeneous accelerator combining NPU and DRAM-based processing-in-memory with hybrid numerical formats for efficient edge LLM inference.
Agentic learner system with grow-and-refine multimodal semantic memory addressing MLLM limitations by retaining domain knowledge across multimodal problem-solving.
Flux4D uses flow-based unsupervised learning for 4D dynamic scene reconstruction without motion annotations, addressing scalability of NeRF and Gaussian Splatting.
Poodle framework for dynamically scaling down LLMs to smaller models through just-in-time replacement for resource-efficient inference on simple tasks.
Phi-table statistical explanation method extending SHAP for tabular model interpretability with directional summaries and uncertainty quantification.
Novel approach using nonequilibrium dynamics in Markov chains for unsupervised generative modeling that spontaneously develops latent-state cycles.
Cross-modal data augmentation framework combining CycleGAN and YOLOv8 for PCB infrared defect detection using unpaired image translation.
LinMU proposes a Vision-Language Model with linear-complexity attention mechanism for efficient edge deployment and high-resolution image/video processing.
Amortized Bayesian inference method for posterior estimation on graph-structured data across diverse domains using permutation-invariant networks.
PAC-Bayesian theoretical framework evaluating performance degradation of neural networks deployed across edge devices with wireless channels.
Bayesian neural networks with singular posteriors via low-rank weight parameterization reducing parameters from O(mn) to O(r(m+n)).
Analysis of attention sink phenomenon in LLMs showing relationship to mixture-of-experts mechanisms with sink-aware training approach.
Ultrafast online learning method using Kolmogorov-Arnold Networks for sub-microsecond adaptation in quantum and fusion control systems.
Analysis of 809 LLM models (2022-2025) quantifies developer-specific efficiency advantages versus compute scaling using regression with fixed effects.
VeRO evaluation harness for assessing coding agent performance on agent optimization through iterative edit-execute-evaluate cycles.
Teacher-student framework using vision-based monocular depth estimation for obstacle avoidance in mobile robot navigation without LiDAR.
Characterizes computational power of recurrent graph neural networks in terms of arithmetic circuits over real numbers.
First-principles theory explaining grokking phenomenon as norm-driven representational phase transition in regularized training dynamics.
Deep reinforcement learning framework using graph neural networks for stage-aware defense against multi-stage advanced persistent threats.
Formulates model retraining for Bayesian prediction systems as cost-sensitive decision problem using posterior learning debt metric based on KL divergence.
Training method unifying supervised fine-tuning with reinforcement learning for LLMs using group advantages and dynamic coefficients.
Adaptive vision foundation model framework for efficient edge deployment via LLM-guided dynamic computation adjustment.
Self-play framework using formal verification to improve LLM code reasoning via semantic equivalence validation in Haskell.
Convergent AI Agent Framework enforcing determinism in agentic workflows via closed-loop control to eliminate constraint violations.
DESPITE benchmark with 12,279 tasks evaluating safety risks of LLM-based robotic planners across physical and normative dangers.
Framework for predicting error patterns in deep neural networks to enable early warning systems for failure.
7-layer neuroscience-inspired memory architecture for autonomous AI systems, outperforming existing memory solutions on long-context evaluation.
Graph-guided loss function incorporating label propagation for fine-tuning language models using global semantic structure.
Mechanistic analysis of why LLM agents deviate from Nash equilibrium in games, with causal intervention experiments on open-source models.
Benchmark for evaluating AI forecasting agents using on-chain validation resistant to overfitting with incentive-compatible scoring mechanisms.
LoRA fine-tuning of 3-4B parameter language models for radiology tasks, enabling CPU deployment in resource-constrained clinical environments.
arXiv paper on Tempus framework for efficient GEMM streaming on AMD Versal for edge LLM inference optimization.
arXiv paper on Themis, multilingual code reward models for flexible multi-criteria scoring in LLM code generation post-training.
Tool for defining business workflows via spreadsheet UI, auto-generating database schema and CRUD APIs using AI.
Historical perspective on Canonical Correlation Analysis (1936) and its relationship to modern JEPA model architectures.
Portable LLM agent framework running in shell script with streaming chat, tool calling, memory, and mentor mode. No dependencies beyond curl and jq.
Product analysis of GPT-5.5's capabilities as coding agent with improved tool use and constraint handling.
ArXiv paper on Group Cognition Learning: multimodal agent collaboration framework addressing modality dominance and spurious coupling in fusion models.
Minimal C runtime (~1000 lines) that executes YAML DAGs where an LLM can edit graph structure mid-execution via four mutation verbs, with full event logging for replay and debugging.
Novel speculative decoding technique using diffusion-style parallel token generation to achieve 3X speedups in LLM inference, replacing sequential autoregressive drafting bottleneck.
File-based context persistence system for AI coding assistants (Claude, Cursor) that preserves project state across sessions without vendor lock-in.
Mtplx: MLX-native inference engine achieving 2.24x speedup on Apple Silicon for LLM token generation. Open source Apache-2.0 license.
Opinion article discussing AI safety practices and risk assessment by AI companies.
Cross-cloud compute provisioning CLI with dataset compression, instance optimization, and 3-5X faster ML deployment with automation pipelines.