CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
CricBench multilingual benchmark for evaluating LLMs on cricket analytics tasks including text-to-SQL and complex statistical reasoning.
CricBench multilingual benchmark for evaluating LLMs on cricket analytics tasks including text-to-SQL and complex statistical reasoning.
Nightjar dynamic speculative decoding adapts draft token verification to varying request loads for improved LLM serving efficiency.
RAIR Chinese benchmark dataset for e-commerce search relevance assessment addressing long-tail and visual salience in LLM evaluation.
FwPKM sparse memory layer balances storage capacity and efficiency for sequence models, addressing softmax attention trade-offs.
STaRR training-free remasking strategy for diffusion language models adapts to spatial-temporal token dynamics for faster inference.
Formal analysis proving LLM recursive self-training undergoes degenerative dynamics without external grounding, limiting self-improvement scaling.
APEX-Agents benchmark evaluates AI agents on long-horizon cross-application tasks in realistic work environments with files and tools.
Theoretical analysis of how rectified flow automatically adapts to intrinsic dimensionality to accelerate sampling in diffusion models.
PyraTok pyramidal tokenizer learns language-aligned discrete video representations across multiple scales for improved video understanding and generation.
Emotion-LLaMAv2 and MMEVerse benchmark for evaluating multimodal LLMs on emotion understanding tasks with high-quality annotations.
Sink token mechanism stabilizes diffusion language models, improving parallel text generation quality and addressing moving sink phenomenon.
TxRay uses agentic approaches for postmortem analysis of blockchain DeFi exploits, analyzing attack patterns and vulnerabilities.
TextME framework projects diverse modalities into LLM embeddings using only text descriptions, enabling multimodal expansion without paired datasets.
EBPO technique using empirical Bayes shrinkage to stabilize Group Relative Policy Optimization for LLM reasoning with verifiable rewards.
Method for merging LLMs across different architectures to transfer knowledge from large models to smaller ones with different designs.
LLM-based reranking approach for recommendation systems using generative reasoning to refine candidate rankings with few-shot learning.
Amortized neural symbolic regression method addressing expression normalization bottleneck to improve scalability for discovering analytical expressions from data.
Systematic characterization of OS-level resource dynamics in sandboxed AI coding agents with resource control mechanisms.
Step 3.5 Flash: sparse MoE model with 11B active parameters for frontier-level agentic reasoning and tool execution.
Action-level off-policy evaluation method for improving vision-language-action models through online RL in robotics.
Study of generative social agents in healthcare motivation dialogues examining effects of knowledge availability on persuasiveness.
Strategic decision framework for governments choosing between buying, building, or hybrid approaches to LLM deployment.
BETA-labeling framework using multiple diverse LLM annotators to construct multilingual low-resource IR datasets.
Framework for self-evolving multi-agent systems to dynamically construct task-adaptive communication topologies for LLM-powered collaboration.
Stability improvement for RL fine-tuning of LLMs by filtering rare spurious tokens to prevent training collapse.
Hardware-aware framework for DNN approximation using multi-level sensitivity scoring and heterogeneous approximate computing blocks.
Using large language models to automatically discover multi-agent reinforcement learning algorithms for imperfect-information games.
Graphic design generation system balancing visual fidelity with structural editability through layered approach.
Transformer architecture analysis for medical time series data addressing temporal and channel dependencies in EEG/ECG.
Unified Memory Agent framework using reinforcement learning to enable LLMs to track state and aggregate evidence in long-context reasoning tasks.
Design-based measurement system using ML-assisted sampling and LLM labeling for cost-effective policy violation prevalence estimation.
Deep reinforcement learning approach for optimal power flow problems in smart grids addressing sample inefficiency.
Proposes stress-gated dynamical regime regulation as alternative to optimization-based learning for autonomous systems operating without fixed objectives.
GIST algorithm for efficient instruction tuning via targeted data selection using coupled optimization geometry and optimizer statistics.
Ensemble method for predicting task affinity in multi-task learning to identify which task groups benefit from joint training.
MapTab benchmark evaluating multimodal LLM reasoning capabilities on constrained route planning tasks with visual grounding requirements.
Diagnostic framework isolating LLM reranker behavior using fixed evidence pools to evaluate ranking policy independent of retrieval quality.
Research on non-interfering weight fields for LLMs to address catastrophic forgetting by treating model parameters as continuously extensible functions rather than fixed weights.
World model using joint-embedding predictive architectures with invariant representations for robust planning.
Time series reasoning system using segment selection and LLMs to adaptively analyze relevant portions of sequences.
Information-theoretic noise schedule allocation for efficient diffusion model training across datasets and resolutions.
Investigates grokking phenomenon showing how transformers develop low-dimensional algorithms with high-dimensional parameters.
Federated learning method for parameter-efficient model adaptation using LoRA merging with theoretical guarantees.
Large causal models framework for temporal causal discovery enabling multi-dataset pretraining.
Studies how transformers learn transfer operators for dynamical systems, enabling zero-shot cross-scale generalization.
Insertion-based sequence generation with learnable order dynamics using discrete flow matching.
Metric distinguishing genuine memorization from generalization in LLMs to assess training data leakage risk.
White-box adversarial attack on world models via physical condition perturbations for autonomous driving systems.
Hierarchical multi-agent reinforcement learning framework for traffic signal control and vehicle eco-driving optimization.
Graph neural network augmented with language models for predicting disease-gene associations in biomedical research.