InfoChess: A Game of Adversarial Inference and a Laboratory for Quantifiable Information Control
Symmetric adversarial game framework for studying information acquisition and inference without material incentives or piece capture.
Symmetric adversarial game framework for studying information acquisition and inference without material incentives or piece capture.
Systematic literature review of LLM-based code summarization techniques for automatic software documentation generation using prompt engineering.
Study on ordered tokenization enabling efficient test-time search in autoregressive generative models through token structure optimization.
Optimized LLM inference kernel for TPU deployment with efficient paged attention mechanism handling dynamic ragged workloads.
Lossless compression approach using chain of lightweight neural predictors for probability estimation in Markov sources.
Study on how supervised fine-tuning increases LLM hallucinations through exposure to new facts and mitigation using continual learning techniques.
Item embedding method for recommender systems that captures temporal dynamics of user preferences beyond bag-of-items approaches.
Proactive AI assistant using LLMs for biomedical discovery, enabling autonomous engagement and foresight in human-AI collaboration workflows.
Attribution analysis comparing interpretive behaviors of LLMs across fine-tuning strategies (FFT, LoRA) for code compliance tasks using perturbation methods.
Adaptive framework for efficient edge deployment of language-aligned vision foundation models with dynamic computation adjustment based on scene context.
Cross-modal retrieval system using multimodal LLM embeddings to match food images with recipe text for cooking and nutrition applications.
Cross-domain benchmark for evaluating LLM self-monitoring and metacognitive abilities across 524 tasks in six cognitive domains using psychometric methodology.
Symbolic reasoning scaffold operationalizing Peirce's abduction-deduction-induction framework with algebraic invariants to improve structured logical reasoning in LLMs.
Framework for training billion-parameter universal ML interatomic potentials without explosive memory growth, enabling quantum-accurate physical simulations.
DiZiNER: instruction refinement method using disagreement-guided pilot annotation simulation to improve zero-shot named entity recognition with LLMs.
RAGognizer: detection head integration method for fine-tuning LLMs to reduce hallucinations and closed-domain inconsistencies in retrieval-augmented generation.
Interpretable machine learning techniques for extracting physical insights from quantum data using variational autoencoders on unlabeled datasets.
SocialGrid benchmark for evaluating LLM agents on planning, task execution, and social reasoning in embodied multi-agent environments inspired by Among Us.
Analysis of output diversity collapse in post-trained language models, showing models produce less varied outputs than base versions, affecting inference-time scaling.
Systematic taxonomy and methods for pruning futile reasoning paths in Large Reasoning Models to improve efficiency and reduce inference costs.
Survey of intrinsic interpretability design principles and architectures for LLMs, focusing on transparency built into models rather than post-hoc explanations.
arXiv paper on constant-factor approximation algorithms for fair k-clustering with demographic constraints.
arXiv paper on backward error analysis and convergence for linear system solvers in numerical linear algebra.
arXiv paper applying YOLOv12 deep learning model for multiclass acute myeloid leukemia cell classification.
arXiv paper on ST-STORM self-supervised learning approach for appearance-based image representation.
arXiv paper on sentiment analysis dataset and LLM-based model for German sign language fairy tales.
arXiv paper on AtManRL method using differentiable attention for faithful chain-of-thought reasoning in LLMs.
arXiv paper on adaptive multi-fidelity optimization with cost-bias tradeoffs and learning rate analysis.
arXiv paper on Information Router to mitigate modality dominance in vision-language models.
arXiv paper proposing framework for informal theorem proving with LLMs using insight-driven reasoning.
arXiv paper on ACSESS method for automatic combination of sample selection strategies in few-shot LLM learning.
arXiv paper proposing HetSheaf, a framework for heterogeneous graph neural networks across different node/edge types.
arXiv paper on HetSheaf framework for heterogeneous graph neural networks supporting multiple node/edge types in real-world applications.
arXiv paper introducing Transformer Neural Processes addressing O(n²) attention bottleneck in Neural Processes with kernel regression.
arXiv paper on Few-Shot Preference Optimization (FSPO) for personalizing LLMs using meta-learning on synthetic preference data.
arXiv paper on AutoNFS, automatic neural feature selection for high-dimensional tabular data with interpretability and efficiency focus.
arXiv paper on Federated Prototype Learning combining textual semantics with visual representations to handle heterogeneous federated learning.
Information-geometric framework for artificial curiosity in sparse-reward RL using intrinsic rewards invariant to representation.
Theoretical and empirical analysis of decentralized learning algorithms comparing multi-stream random walk and asynchronous gossip approaches.
Histogram-based parameter-efficient tuning (HPT) technique for transfer learning capturing target domain statistics in sonar classification.
Softpick: Rectified softmax replacement for transformer attention eliminating attention sink and reducing activations with improved quantization.
ChemAmp: Framework for composable LLM agents in chemistry using tool amplification to enhance multi-tool orchestration capabilities.
Token significance-aware RL method for LLM reasoning that optimizes token-level contributions rather than uniform length penalties.
PyLO: PyTorch package making learned optimizers accessible, providing drop-in replacements for standard optimizers like Adam.
HiPreNets: Progressive training approach for neural networks achieving high precision in L-infinity norm error for safety-critical applications.
Analysis of exploration-exploitation bias in offline evaluation of linear bandit recommender systems using contextual bandits.
Self-aligned reward (SAR) method for LLM reasoning that provides fine-grained guidance beyond binary correctness to improve efficiency and accuracy.
Distributionally robust optimization approach for RLHF alignment addressing overoptimization in LLM training via relative reward regression.
SmilesGEN: VAE-based generative model using multi-objective RL for de novo drug molecule generation considering phenotypic effects.
Empirical study of scaling behaviors in RL post-training for LLMs across Qwen2.5 models (0.5B-72B) focused on mathematical reasoning performance.