AIvaluateXR: An Evaluation Framework for on-Device AI in XR with Benchmarking Results
AIvaluateXR: evaluation framework for benchmarking LLMs on XR devices, testing 17 models for on-device inference performance and selection.
AIvaluateXR: evaluation framework for benchmarking LLMs on XR devices, testing 17 models for on-device inference performance and selection.
DeePen: penetration testing methodology for evaluating robustness of machine learning audio deepfake detection classifiers.
Offline reinforcement learning method using value function inconsistency penalties to improve model-based policy learning from fixed datasets.
HiddenBench: 65-task benchmark evaluating multi-agent LLM systems' ability to reason with distributed information in collective settings.
Region-based reinforcement learning approach optimizing LLM performance for table question-answering with structured row-column reasoning.
A3: analytical low-rank approximation framework for transformer attention mechanisms to enable efficient LLM compression and deployment.
Framework combining algorithmic information theory and neural network pruning to improve model generalization through data compression.
Analysis of vision tokens in large vision-language models to reduce hallucinations and improve multimodal decoding with visual semantic guidance.
Dilated Unmasking Scheduler for fast non-autoregressive text generation in masked diffusion language models with improved parallelization.
LoRA-Mixer: modular mixture-of-experts framework routing task-specific LoRA experts through attention matrices for efficient multi-task LLM adaptation.
RIGVid: system enabling robots to learn complex manipulation tasks by imitating AI-generated videos without physical demonstrations.
Agent-driven system for mining rare disease information from unstructured clinical notes using LLMs to improve medical documentation.
Analytical framework modeling autoregressive language models as composition of information-processing stages using Markov categories.
Exact verification method for graph neural networks using incremental constraint solving to provide adversarial robustness guarantees.
GLASS: inference-time sparsification method for LLMs using global-local aggregation to improve deployment on resource-constrained devices.
Fine-tuned LLaMA 3.2 vision-language model adapted for neutrino event classification in high-energy physics detector data.
Diffusion-based world models for offline reinforcement learning that generate actions alongside states and rewards for improved TD learning compatibility.
Theoretical study of linear models for time series forecasting, analyzing their robustness and interpretability through characteristic root analysis.
Security method detecting impersonation in AI videoconferencing by exploiting biometric leakage from pose-expression latents to expose puppeteering attacks.
Framework for certifying tool selection in LLM-based agentic systems, evaluating robustness against adversarial tool pools and deployment scenarios.
Robot policy learning method combining continuous and discrete representations using vision-language and visual dynamics models for diverse manipulation tasks.
Study comparing expert pruning vs merging strategies for compressing Mixture-of-Experts models, finding pruning superior for generative tasks.
Automated safety evaluation framework for mental health AI chatbots using clinician-informed rubric and multi-agent validation approach.
Latent-augmented discrete diffusion model with auxiliary channel for improved few-step language generation and cross-token dependencies.
Study on trade-offs between model accuracy and inference efficiency examining architectural factors affecting LLM deployment costs.
Theoretical analysis of sample complexity in differentially private policy optimization for reinforcement learning with privacy guarantees.
Benchmark evaluating whether LLMs can iteratively develop code toward high-level goals beyond isolated task completion.
Sequential autoregressive generation with reinforcement learning and MCTS to enforce hard constraints in planning and design tasks.
Benchmark for evaluating multimodal LLMs on streaming video understanding with human gaze signals for AR applications.
Memory-efficient LLM training technique using blocked state folding to reduce memory bottlenecks with Adam optimizer.
Theoretical analysis of feature evolution dynamics in infinite-depth ResNets under depth-μP scaling during training.
Deep Delta Learning introduces residual update rule for transformer layers enabling selective content rewriting in deep networks.
Multi-agent framework using retrieval-augmented generation and LLM judges to create automated rubrics for evaluating medical dialogue systems.
Research on quantized matrix multiplication optimization for efficient LLM deployment with weight and activation quantization.
Benchmark evaluating privacy risks in GUI agents that process screenshots, measuring exposure of sensitive information during operation.
Framework using LLMs to generate synthetic training data for biomedical entity linking, addressing scarcity of expert annotations.
Principled optimization-based interpretation of classifier-free guidance in flow matching models for controllable generation tasks.
Theoretical analysis of gradient descent training dynamics, generalization bounds, and differential privacy properties for Kolmogorov-Arnold Networks (KANs).
RAG-GNN framework combines graph neural networks with retrieved biomedical literature for cancer signaling predictions using contrastive alignment and gated fusion mechanisms.
Post-training quantization method balancing rank budgets between preserving low-rank structure and reconstructing quantization error in LLMs.
Theoretical analysis of learnable Bernstein polynomial activations showing exponential approximation rates and parameter efficiency.
Entropy-aware reward guidance mechanism for test-time adaptation in discrete diffusion language models.
Single-trajectory verification framework that steers reasoning models by providing intermediate feedback during inference.
Analysis revealing test-time training with KV binding is equivalent to learned linear attention, challenging memorization interpretation.
Empirical study showing LLMs differ from humans in goal selection during self-directed learning, challenging their use as proxies for human preferences.
Benchmark and framework for multi-party negotiation games with binding action-level commitments derived from real climate negotiation data.
Optimized sampling primitive that fuses categorical sampling into LM-head computation, reducing memory traffic for large-vocabulary decoding.
LLM-driven algorithmic debugging approach for abstract reasoning tasks using abduction-based program refinement.
Distributed reinforcement learning method addressing negative learning from high-surprisal data in stale actor settings.
Spoken dialogue model with controllable response duration for voice assistants and interactive agents.