MobileKernelBench: Can LLMs Write Efficient Kernels for Mobile Devices?
MobileKernelBench benchmark evaluating LLM capability to generate efficient computational kernels for mobile devices, with systematic investigation of code generation limits.
MobileKernelBench benchmark evaluating LLM capability to generate efficient computational kernels for mobile devices, with systematic investigation of code generation limits.
LoV3D: Vision-language model pipeline for longitudinal brain MRI analysis that grounds neurological disease progression reasoning in regional volume measurements.
Research on model stitching technique for Vision Foundation Models, testing representational compatibility across models with different training objectives and data sources.
Method for incorporating text into time-series forecasting by bridging the modality gap between qualitative text and quantitative forecasting signals through semantic space alignment.
MetaKE addresses knowledge editing in LLMs using bi-level optimization to fix specific facts without degrading general capabilities, identifying semantic-execution misalignment issues.
Empirical study of model collapse in large language models trained recursively on synthetic data.
Self-evolution framework using Minimum Bayes Risk decoding for error span detection in machine translation without human annotations.
Large-scale open-source software engineering environment for training AI agents with executable, verifiable tasks and dynamic feedback.
Method for continual fine-tuning of pre-trained models on sequential tasks with parameter-free task retrieval and no forgetting.
Code agent framework with structured memory enabling adaptive learning from project evolution and past successful reasoning trajectories.
World model architecture using spherical kernel operators to handle shifting data distributions in latent space transitions.
Federated framework combining lightweight LLMs with personal knowledge graphs for privacy-preserving personalized recommendations.
Neuro-symbolic architecture combining self-supervised learning with verifiable logic rules to mitigate spurious correlations and shortcut learning.
Self-distillation method reducing computational cost of chain-of-thought reasoning by training models to generate correct predictions from truncated reasoning.
Zero-shot LLM approach for surgical duration prediction combining retrieval-augmentation with Bayesian averaging, avoiding need for fine-tuning.
Tree-based continual learning framework for non-stationary data distributions with constrained computational resources in time series applications.
Sparse autoencoders foundation for learned sparse retrieval, decomposing LLM representations into interpretable latent features for efficient document retrieval.
Multi-model inference optimization reusing identical KV caches across models to reduce memory consumption in agentic AI systems.
Federated learning method for LoRA fine-tuning of LLMs addressing statistical and functional heterogeneity across model layers.
Benchmark quantifying LLM robustness by measuring model sensitivity to prompt variations, typos, and paraphrases in real-world conditions.
Framework reframing LLMs as code generators for interpretable decision-making in high-stakes scenarios, improving reproducibility over black-box approaches.
KV cache optimization technique for multi-agent LLM systems that reuses decoding caches to reduce memory usage and latency in collaborative AI tasks.
Federated learning framework for multimodal sentiment analysis using uncertainty-aware fusion to handle missing modalities and heterogeneous data.
Pragma-VL approach balancing safety and helpfulness in multimodal LLMs through pragmatic alignment methods.
ICPRL framework enabling vision language models to learn physical reasoning from pixel-based interactive control.
DreamReader: unified interpretability toolkit for analyzing text-to-image diffusion models with causal and representation analysis.
Two-stage ML framework for identifying high-performance nested antiresonance fiber designs in telecommunications.
Approach for stable distributional alignment in LLMs to predict population response distributions across options.
Framework for training single models to answer conditional queries across heterogeneous datasets via task expansion.
Method addressing curriculum collapse in self-improving LLM reasoning systems through diverse problem generation.
Neural basis functions using untrained networks for multivariate function approximation in machine learning.
Research on linear structure in LLM attention head activations for KV cache optimization in Transformer inference.
Evaluates LLMs for gait classification from text-encoded kinematic waveforms, comparing performance to ML methods for clinical interpretability.
Analyzes overfitting and false refusals in fine-tuned LLMs through residual stream analysis, identifying safety data entropy as key factor.
LightningRL uses reinforcement learning to improve block-wise diffusion LLMs, breaking the accuracy-parallelism trade-off in parallel token generation.
Modular Neural Computer is a memory-augmented architecture combining external associative memory with functional MLP modules for exact algorithmic computation.
Studies out-of-distribution detection in motor imagery brain-computer interfaces to prevent misclassification on unseen data distributions.
Proposes feature-level interaction explanations for multimodal transformers, identifying cross-modal synergy and redundancy in predictions.
RBF-Solver proposes radial basis function-based multistep sampling for diffusion models to accelerate inference without predefined schemes.
AdaBox introduces adaptive density-based clustering with parameter generalization to reduce hyperparameter sensitivity across datasets.
Framework integrating vehicle sensor streams with contextual signals for V2X-augmented predictive maintenance using multi-dataset evaluation.
PolyGLU enables transformers' feed-forward neurons to dynamically route among multiple activation functions, improving architectural flexibility.
Real-time conversational AI system combining speaker segmentation with hierarchical end-of-turn detection for natural two-speaker voice interactions.
Proposes GPrune-LLM, a structured pruning method for LLMs that estimates neuron importance using distribution-robust techniques for better cross-task generalization.
Analyzes how diffusion models generalize despite optimal models memorizing training data, showing denoising trajectory properties affect generalization.
Investigates memorization behaviors in Rectified Flow generative models for image synthesis through theoretical and empirical analysis.
Proposes Kalman World Models, a method for training state-space models using recursive Bayesian filtering instead of backpropagation for online learning.
Flow matching-based approach for full-waveform inversion with generative priors to improve seismic imaging robustness.
Outcome-Aware Tool Selection method for semantic routers in LLM inference, reducing latency by offline interpolation without GPU cost.
Standardized benchmark dataset and evaluation framework for computational antibody design methods with unified metrics.