EcoFair: privacy-preserving medical inference framework with lightweight routing for vertically partitioned data and modality-specific embeddings.
Theoretical analysis of spectral optimizers like Muon in language model training, studying capacity scaling through linear associative memory framework.
Machine unlearning framework addressing retain-forget entanglement where retained samples unintentionally affected by forgetting correlated features.
Study comparing sample selection methods (random, farthest-first, interactive visualization) for annotation of biomedical time-series data with real annotators.
PQuantML: open-source hardware-aware neural network compression library for pruning and quantization with unified interface for latency-constrained deployment.
Quantum-inspired anomaly detection using hardware-aware tensor networks for particle collider physics, deployable on classical hardware.
Systematic evaluation of tabular foundation models like TabPFN and TabICL for conditional density estimation in regression tasks with heteroscedasticity.
C²MF: context-specific credibility-aware multimodal fusion framework using probabilistic circuits to handle conflicting modalities and situational reliability changes.
Judge Agent system using automated mathematical validation to reduce silent failures in LLM-generated scientific simulation code from 42% to 1.5%.
LLM framework for formal proof repair using counterexample-guided reasoning and behavioral feedback to improve automated verification.
Comprehensive evaluation framework for agent-based medical AI systems via multi-step clinical dialogue simulation with realistic physician-patient interactions.
Real-world evaluation of visual navigation foundation models on robot navigation, testing generalization and providing trajectory quality metrics.
Vision-language learning approach for end-to-end autonomous driving using multimodal datasets and collision-aware representation learning.
Optimization framework for robust decisions when predictions lack calibrated error bounds, combining robust and regret formulations.
Theoretical analysis of rank selection for low-rank tensor regression with applications to neural network compression and model optimization.
Decoupled audio transformer architecture inspired by human cognition for efficient self-supervised learning on resource-constrained devices.
Analysis of object discovery in self-supervised Vision Transformers, showing how [CLS] token attention maps contain spurious activations affecting localization.
Continual learning method for object detection under extreme visual sparsity conditions using dual-stage invariant learning.
Higher-order associative memory models combining exponential interactions with sparse pattern storage for improved storage capacity.
Analysis of privacy-accuracy trade-offs in high-dimensional sparse linear regression using differential privacy mechanisms and approximate message passing.
Method for compressing conversational audio context in LLM-based speech recognition systems, studying multimodal context from prior turns for improved ASR.
Approach for merging multiple LoRA modules while preserving subspace coverage and addressing directional anisotropy to maintain task representation in general-purpose systems.
Benchmark for evaluating machine unlearning in multimodal models like CLIP, introducing SALMUBench with 60K persona-attribute associations for fine-grained forgetting evaluation.
Method for merging independently fine-tuned LoRA adapters across heterogeneous tasks using null-space compression, addressing classification-regression task combinations.
Graph-learning algorithm (MED-MAGMA) for fitting Kronecker-sum-structured models with multiplicative noise in genomics applications.
Generative approach for uncertainty quantification in multimodal supervised learning combining images and text data.
Theoretical analysis of Kantorovich-kernel neural network operators with density results, convergence estimates, and Korovkin theorems.
Meta-learning framework for human mesh recovery from images using optimization-friendly initializations and uncertainty-aware updates.
UNIFERENCE: discrete-event simulation framework for developing and benchmarking distributed AI inference algorithms across heterogeneous devices and networks.
AMALIA: fully open source LLM trained on high-quality European Portuguese data with native evaluation benchmark and improved pt-PT representation.
ALBA: linguistically grounded benchmark for evaluating LLM performance on European Portuguese, addressing underrepresentation in existing benchmarks.
Experimental pipeline profiling energy consumption, latency, and quality trade-offs for deploying LLMs on edge devices with hardware constraints.
Study evaluating ML feature compatibility and transferability across malware detection datasets under distribution shifts.
Deep symbolic regression using policy gradients with complexity awareness for interpretable data-driven mathematical expression discovery.
Theoretical analysis of how iteration order affects convergence and stability in deep neural network training without learning rate schedules.
Methodological commentary on robust predictive modeling under distribution shifts in real-world deployment scenarios.
Task Tokens method adapts behavior foundation models to specific tasks via learnable tokens while preserving zero-shot generalization capabilities.
FastCache accelerates Diffusion Transformer inference through learnable linear approximation and spatial-aware token selection for hidden-state caching.
Defends RAG systems against knowledge poisoning attacks by detecting and mitigating adversarial text injections in external knowledge sources.
PepThink-R1 integrates LLMs with chain-of-thought supervised fine-tuning and reinforcement learning for interpretable cyclic peptide design optimization.
LLMs perform automatic wireless modulation classification via discretized self-supervised candidate retrieval, avoiding distribution shift issues of supervised models.
Control-theoretic framework for LLM activation steering with feedback controllers, connecting empirical steering methods to proportional control theory for safety alignment.
NeST-BO proposes Newton-step targeting Bayesian optimization using Gaussian processes to learn gradient and Hessian information for expensive black-box problems.
Sequence-level TopK (SeqTopK) improves Mixture-of-Experts routing in LLMs by adapting expert assignment per sequence rather than per token without retraining.
Cascading Bandits analyzes decision-making policies for edge inference with multiple models, providing theoretical regret guarantees for Explore-then-Commit and Thompson Sampling approaches.
LiteCache optimizes KVCache memory management for LLM inference using GPU-centric query similarity-driven approach to reduce memory overhead and improve CUDA Graph execution.
Repulsive Bayesian Prompt Learning addresses overfitting in prompt learning for foundation models using Bayesian inference framework for improved out-of-distribution generalization.
Balanced Fine-Tuning aligns LLMs with biomedical knowledge through confidence-weighted token-level optimization and adaptive reward mechanisms.
FedRE proposes a representation entanglement framework enabling federated learning across clients with heterogeneous model architectures and data.
SonicMoE optimizes Mixture of Experts inference through IO-aware and tile-aware techniques for high-granularity, sparse MoE language models.