CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
Dataset and ML models for transpiling GPU code between CUDA and HIP architectures, with 60k verified code pairs.
Dataset and ML models for transpiling GPU code between CUDA and HIP architectures, with 60k verified code pairs.
Hierarchical RL approach using multi-resolution skills to improve manager subgoal selection and performance on agile tasks.
Novel multi-Boolean architecture framework for binarizing LLM weights during training for efficiency without full-precision latents.
LLM-based data augmentation framework for recommendation systems using majority-voting reranking to address data sparsity.
Reinforcement learning agent for quantitative trading combining multi-indicator technical analysis with RL for adaptive decision-making.
LoRA-based method for efficient continuous concept control in diffusion models for image/video synthesis with plug-and-play design.
Research on sequential fine-tuning of large multimodal models showing skill recovery across different tasks and model families.
Paper on efficient autoregressive inference for transformer probabilistic models balancing set-conditioning with joint distributions.
Paper on test-time prior adaptation for simulation-based inference using diffusion models for Bayesian inference.
TROJail uses trajectory-level optimization with process rewards to learn multi-turn jailbreak strategies against LLMs.
Stream.FM applies flow matching for real-time streamable speech restoration with 32ms algorithmic latency.
DRAM framework combines mechanism design and online learning for sequential multi-agent truthful reporting.
Fitted Q-evaluation theory for off-policy reinforcement learning without requiring Bellman completeness using stationary weighting.
QSLM quantization framework with tiered search optimizes spike-driven language models for embedded deployment.
Study investigates whether LLMs encode functional importance of individual reasoning tokens for reasoning chain compression.
Theoretical analysis of local updates in distributed optimization showing acceleration benefits and topology effects in federated settings.
SAGE-32B is a 32B parameter model fine-tuned via iterative distillation for agentic reasoning, task decomposition, and tool usage.
Multi-Focus Attention Instruction probe disentangles recognition vs synthesis failures in multi-hop LLM reasoning.
Study reveals LLMs are highly sensitive to prompt order in multiple-choice QA due to causal attention limitations.
Temp-R1 is an autonomous agent for temporal knowledge graph question answering trained via reverse curriculum reinforcement learning.
AskBench evaluates and improves LLM ability to request clarification on ambiguous prompts using reinforcement learning with rubric guidance.
CLIPoint3D adapts vision-language models like CLIP for 3D point cloud domain adaptation with few-shot unsupervised learning.
ConFu improves speculative decoding for LLM inference acceleration by enhancing draft model quality to propose better candidate tokens for verification.
Self-distillation degrades LLM reasoning by suppressing epistemic verbalization of uncertainty; controlled experiments isolate degradation mechanisms.
PolarQuant post-training quantization for LLMs uses Hadamard rotation and Gaussian weight distribution for near-lossless compression.
Analysis reveals multilingual language models organize internal representations by orthographic script rather than linguistic structure across language families.
Semantic Intent Fragmentation attack exploits LLM orchestration systems where composed subtasks violate security policy despite individual benignness.
Triadic Suffix Tokenization improves LLM numerical reasoning by partitioning digits into three-digit triads with explicit magnitude markers.
Multi-agent study evaluating whether LLMs can cooperate on resource governance through elected leadership and self-governance mechanisms.
Event Tensor abstraction eliminates kernel launch overheads in LLM inference by fusing operators into persistent kernels handling dynamic shapes.
Cross-domain metacognitive benchmark for LLMs with 524 items across six cognitive domains using human psychometric methodology.
BARD framework bridges autoregressive and diffusion vision-language models via progressive block merging and stage-wise distillation for efficient inference.
Q-SINDy integrates quantum feature maps into sparse identification of nonlinear dynamics, addressing coefficient cannibalization failure mode.
Novel framework integrating Uniform Discrete Diffusion Models with Group Relative Policy Optimization for stable RL training on discrete generative models.
Systematic benchmark comparing cloud and open-source LLMs on System Dynamics tasks: causal loop diagram extraction and interactive coaching.
Incomplete article about arXiv Labs framework; content does not match title about LLM benchmarking.
Aide: customizable Android voice assistant supporting Claude, OpenAI, or OpenAI-compatible endpoints with on-device encryption.
Open source local screen memory tool for Claude and coding agents; OCR and summarization via local AI; Mac-only Swift app.
Prismer: infrastructure layer for long-running AI agents with error recovery, persistent memory, and cross-session learning.
Cloudflare's internal AI engineering stack: MCP servers, agent infrastructure, and iMARS tiger team integration for engineering workflow.
Marketing copy for Meticulous automated testing tool for AI-generated code; claims to eliminate debugging.
MemFactory: unified framework for training and inference of memory-augmented LLMs using reinforcement learning for agent memory operations like extraction and retrieval.
MemFactory: unified framework for inference and training of memory-augmented LLMs for long-term AI agents using RL optimization.
Technical article on building search engines for AI agents, handling complex query graphs and agent-generated code constraints.
ASCEND: DevSecOps framework with AI-powered merge conflict resolution integrated into CI/CD pipelines.
Meta installing tracking software (MCI) on employee computers to capture interactions for AI agent training.
Kuri: Zig-based browser automation tool for AI agents with 464KB binary, 3ms cold start, 16% token efficiency improvement.
Benchmark study measuring vulnerability patterns in code generated by AI models under time pressure conditions.
Google WeatherNext 2: state-of-the-art ML models for weather forecasting released to researchers and enterprises.
PayClaw: tool enabling AI agents to manage wallets and execute financial transactions autonomously.