arXiv paper introducing KGroups, a feature selection algorithm for high-dimensional biological data using max-relevance min-redundancy criteria.
IsoQuant uses quaternion algebra and isoclinic rotations for efficient LLM KV cache compression with hardware-aligned blockwise operations.
FeDMRA addresses federated class-incremental learning with dynamic memory replay for non-IID distributed healthcare data.
HISA improves efficiency of token-level sparse attention mechanisms through hierarchical indexing, reducing O(L²) bottleneck.
Analysis of scaling laws in AI across model families, explaining their predictive power and universal effectiveness in training loss reduction.
CirrusBench evaluates LLM-based agents in real-world cloud service environments beyond correctness, measuring robustness and efficiency.
Simplex denoising framework for discrete generative modeling using non-Markovian noising scheme, applied to graph generation.
Offline multi-agent reinforcement learning approach using Partial Action Replacement to handle exponential joint action space growth.
ChemCLIP uses contrastive learning to bridge organic and inorganic anticancer compound discovery by enabling knowledge transfer across chemical domains.
LACE mechanism for continual learning that adaptively expands model capacity during training based on loss signal monitoring.
Information-theoretic analysis of safety verification impossibility for self-improving systems balancing bounded risk with unbounded utility.
AMIGO benchmark for evaluating agentic vision-language models on long-horizon multi-image grounding tasks through sequential attribute-focused queries.
GPU-accelerated TensorRT inference pipeline for BERT and GPT-2 with mixed-precision optimization achieving 64.4x CPU speedup.
VeoPlace uses vision-language models for chip floorplanning macro placement by leveraging VLM spatial reasoning abilities to complement learning-based approaches.
HyperP introduces hypersphere parameterization for language model scaling with improved training stability compared to first-order optimizer approaches.
Analysis of why linear probes and sparse autoencoders fail at compositional generalization under superposition, proposing iterative coding alternatives.
SimulCost benchmark evaluates LLM agents on physics simulation tasks with cost-aware metrics, accounting for simulation time and experimental resource usage beyond token costs.
Conversational query rewriting approach for multimodal image retrieval with multi-turn dialogue dataset.
GeoBlock infers optimal block sizes for diffusion language models by analyzing token dependency geometry to enable efficient parallel decoding.
FEMBA: bidirectional Mamba state-space model pre-trained on 21k hours EEG with physiologically-aware objectives for microcontroller deployment.
Enhanced mixture-of-experts architecture using soft nearest neighbor loss to prevent expert collapse and redundant representations.
Cross-lingual evaluation of vision-language models on visual reasoning tasks across Indian languages, revealing performance disparities.
Framework integrating sparse autoencoders with dynamic head pruning in Vision Transformers for interpretable and controllable efficiency.
MotionGPT3 replaces diffusion with rectified flow objectives for efficient text-driven motion generation with improved convergence.
Evolution strategies warm-start reinforcement learning agents for industrial continuous control using CMA-ES-generated demonstrations.
LogicDiff improves reasoning in masked diffusion language models by prioritizing logical connective tokens during inference-time denoising.
Privacy-preserving inference for spiking neural networks using fully homomorphic encryption to enable encrypted computation.
Language-conditioned multi-game procedural level generation using shared neural representations across different game domains.
Study of why minimal GPTs fail at out-of-distribution arithmetic generalization, revealing staged failures from layout barriers to positional encoding limits.
Survey of uncertainty-aware explainable AI methods, examining integration of uncertainty quantification (Bayesian, Monte Carlo, Conformal) into explanatory pipelines.
Controlled evaluation of LLM implementation choices for political text annotation, testing model selection, size, and prompt engineering best practices.
Comparative evaluation of physics-informed neural networks and neural ODEs for modeling nonlinear neuronal dynamics on Morris-Lecar model.
KOMET: model-agnostic framework using Koopman operators to track parameter evolution and handle temporal domain drift in non-stationary environments.
ASTER: agentic toolkit using LLMs for exoplanet research workflows, combining archival queries, literature search, and radiative transfer models.
Online statistical inference framework for sample-averaged Q-learning to reduce variance and improve stability in reinforcement learning.
Analysis of reliability limits in LLM-based multi-agent planning systems modeled as decision networks with language-based communication constraints.
FormalProofBench: benchmark evaluating whether LLMs can produce formally verified graduate-level mathematical proofs using Lean 4.
Lightweight neural network for super-resolution imaging from low-resolution SPAD arrays, reconstructing 256x256 images on embedded devices.
Comparative study evaluating YOLO object detection models on robotics tasks using custom and COCO2017 datasets for workspace object detection.
Study of incentive collapse paradox in AI-assisted task delegation showing accuracy improvements require unbounded payments without intervention mechanisms.
Persona-based LLM approach for simulating diverse human opinions at population scale for social science interventions and consequence modeling.
Theoretical analysis of loss landscape geometry in regularized deep matrix factorization proving unique minimizers under weight decay.
Information-theoretic framework for forecasting measuring mutual information between future observations and information set as predictability limit.
Sovereign Context Protocol defines open runtime attribution layer for human-generated content used in LLM training and inference.
Systematic evaluation of segmentation and geospatial foundation models for global field boundary segmentation using FTW benchmark.
Bayes-MICE extends multiple imputation for time series missing data using Bayesian inference and MCMC sampling.
Conformal Prediction Assessment framework for evaluating conditional coverage validity in distribution-free prediction with finite-sample guarantees.
StretchCast global-regional AI weather forecasting framework using variable-resolution cubed-sphere mesh for refined regional predictions.
Comparative analysis of AI datasets, foundation models, and barriers to achieving general-purpose AI in surgical image analysis.
D-SPEAR dual-stream replay mechanism for stable off-policy reinforcement learning in robotic manipulation with contact-rich dynamics.