Difficulty-Differentiated Policy Optimization addresses Large Reasoning Models' overthinking and overconfidence by redistributing token allocation based on problem difficulty.
Framework for assessing information security awareness in LLMs, including security knowledge, attitudes, and behavior to improve rejection of unsafe requests.
Genomic language model framework using phylogenetic trees and multispecies alignment for identifying evolutionarily constrained sequences.
Analysis of multi-stage LLM inference pipelines including RAG, KV cache retrieval, routing, and reasoning with optimization strategies.
Pseudo-simulation method for evaluating autonomous vehicles addressing limitations of real-world and closed-loop simulation evaluation.
Framework for multimodal representation learning through simultaneous alignment of diverse data modalities.
Method leveraging superclasses for representation disentanglement to mitigate spurious correlations and improve group robustness.
Evaluation-Aware RL framework considers policy evaluation accuracy during training to reduce variance and bias.
Method for detecting intersectional bias in face recognition embeddings using directional alignment in latent space.
CARES lightweight module selects appropriate image resolution for vision-language models to reduce token overhead and latency.
RobotArena∞ enables scalable robot benchmarking through real-to-sim translation for evaluating diverse robotic agents.
Rep2Text framework recovers original input text from single LLM token representation using trainable adapter for interpretability.
FastMMoE accelerates multimodal LLM inference through dynamic expert activation and token pruning for reduced latency.
Unsupervised feature selection method using robust autoencoder and adaptive graph learning for high-dimensional data clustering.
Dementia-R1 applies reinforced pretraining and reasoning to LLMs for longitudinal clinical prognosis from unstructured medical notes.
Benchmark and moderation model for evaluating LLM safety, adversarial robustness, and handling of nuanced harmful content detection.
Case study integrating PubChem, ChEMBL, and eMolecules using byte-offset indexing for terabyte-scale chemical database. Infrastructure for ML-driven molecular property prediction.
Framework for managing ambiguity in long-horizon workflow agents. Task-agnostic approach for curating and measuring impact of underspecified instructions on agent execution.
Method using diffusion models to enhance CLIP visual representations by improving both discriminative ability and fine-grained detail perception.
Study of vision language models for spatial grounding in 3D medical imaging. Examines VLM performance across imaging modalities and slice directions.
AC-Foley framework for video-to-audio synthesis using reference audio guidance and acoustic transfer. Addresses semantic granularity and acoustic feature description challenges.
Security research on ClawWorm, self-propagating attacks across multi-agent LLM ecosystems. First study of attack propagation in interconnected agent systems like OpenClaw.
Theoretical analysis of partial label learning feasibility and adaptive nearest neighbor methods. Mathematical characterization of PLL learning conditions.
IRIS benchmark with 220 high-fidelity 4K videos for physical parameter estimation and governing equation identification from monocular video.
Research using Stochastic Gumbel AlphaZero to evaluate game difficulty in Tetris Block Puzzle variants. Applies game-playing AI as evaluation metric.
Multi-agent AI orchestration for modernizing legacy COBOL banking systems using Claude MCP. Standards-compliant AI agent architecture.
iOS SDK for embedding AI agents with tool calling, memory, and thread management. Production-grade AI agent framework for mobile.
Firefox extension blocking procrastination on YouTube/Reddit, built with Claude Code. LLM application but minimal technical detail.
Multi-agent coding assistant with LangGraph orchestration and sandboxed Rust execution engine. Production AI agent framework.
Research showing language models can de-anonymize forum posts using writing style analysis. LLM capability study.
AgentVerse is an open-source social network platform for AI agents to interact and collaborate.
Curated tool setup for Claude Code-based development with 15 opinionated tools. Practical guide for AI-assisted coding workflows.
Proposes 'Franny Test,' a three-step adversarial protocol to expose structural limitations in LLM reasoning and imitation capabilities.
Cursor's Composer 2 coding model revealed to be based on Moonshot AI's open-source Kimi 2.5 with additional fine-tuning.
Research paper analyzing self-recursive ethics in AI systems using a 6,334-entry ethics monitor log spanning seven months.
Promotes proprietary software claiming to be a transformer alternative (Mamba, Hyena, RWKV competitors) with C/C# implementation.
Discussion on developer spending for AI coding tools like Cursor and Claude Code at work.
Open-source Claude skills for Git workflow automation and weekly summary generation.
Discussion on tools and methods for comparing cloud and AI costs across providers.
Analysis of how agentic AI systems generate massive token usage and costs that exceed traditional per-token pricing models.
Research proposing continuous vector prediction instead of token prediction for LLMs. Novel architecture improving efficiency and reasoning.
Plexus is an API gateway unifying access to multiple LLM providers (OpenAI, Anthropic, Google, etc.) under a single endpoint, allowing model/provider switching without code changes.
AskAlf: Self-hosted AI agent platform that creates specialized AI workers for tasks (marketing, support, research). Runs locally on Docker, 24/7 operation.
Novel technique applying video compression principles to LLM KV cache during inference, achieving 10,000x less quantization error at same storage cost.
Brief reference to Claude Code for academic use. Incomplete PDF content, minimal information.
MCP tool for AI agents/LLM tools (Cursor, Claude) that sends phone notifications when long-running tasks complete, reducing context-switching.
Walmart-OpenAI ecommerce partnership through ChatGPT shows disappointing sales, suggesting AI agent adoption in commerce slower than expected.
CapKit: 200-line open-source library providing scoped, time-bound, cryptographically-signed capabilities for AI agents to limit permissions and prevent privilege escalation attacks.
Analysis of local open-source AI models as alternative to datacenter-dependent systems. Discusses performance parity with frontier models within 6 months.
Open-source note-taking application alternative to Obsidian featuring MCP server integration for AI capabilities.