Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
Lil: Study showing sparse-attention algorithms underperform in LLM decode stage despite improvements in prefill stage.
Lil: Study showing sparse-attention algorithms underperform in LLM decode stage despite improvements in prefill stage.
Analysis of whether LLMs internally encode token-level functional importance in reasoning chains for compact reasoning generation.
MMErroR: 1997-sample benchmark evaluating whether Vision-Language Models detect erroneous reasoning across 24 subdomains.
BiForget framework for automated synthesis of high-quality forget sets addressing domain and instance-level unlearning granularities in LLMs.
Tape: Cellular automata benchmark isolating rule-shift generalization in RL with controlled protocol and continuous metrics.
Compositional steering tokens enable multi-behavior control of LLMs simultaneously, addressing underexplored problem of steering toward multiple objectives.
GanitLLM: Bengali mathematical reasoning model with difficulty-aware corpus and curriculum-based GRPO training pipeline for multi-step math problems.
Training-free method to mitigate hallucinations in vision-language models by using attention-space contrastive guidance for visually grounded generation.
Systematic analysis of demographic bias in LLM-generated targeted messaging across GPT-4o, Llama-3.3, and Mistral models with evaluation framework.
Benchmark for legal fact-checking: evaluates systems on verifying layperson legal claims against U.S. Supreme Court precedents.
Benchmark for evaluating vision-language model reasoning: assesses whether VLMs can explain geolocation predictions with supporting image evidence.
SIGMA multi-task recommender system using LLMs for semantic-grounded instruction-driven generative recommendations in e-commerce.
ReflexiCoder teaches LLMs to self-reflect and self-correct generated code via reinforcement learning without external oracles.
Design-time verification framework for trustworthy AI enabling pre-training verification of model correctness and stability.
LogicDiff improves zero-shot reasoning in masked diffusion language models via logic-guided token unmasking at inference time.
SkillX automatically constructs plug-and-play skill knowledge bases for LLM agents to improve learning efficiency and generalization.
ConsistRM improves generative reward models via consistency-aware self-training to address human preference alignment at scale.
LLM-based framework for validating and restructuring unsupervised text clusters using semantic reasoning and refinement.
Self-evolving agents framework jointly optimizing policy and tool graph memory for autonomous learning and tool synthesis.
LLMs for clinical decision support with transparent reasoning using Toulmin argumentation and curriculum learning to improve trustworthiness.
ProbeLogits presents kernel-level LLM inference operations for AI-native OS, enabling logit-based classification of agent actions without learned parameters.
VoxSafeBench introduces a benchmark for evaluating safety in speech language models within multi-user environments, considering speaker identity and environmental context.
VeriGraphi proposes a multi-agent LLM framework for generating hierarchical Verilog hardware designs, addressing context loss and hallucination issues in RTL generation.
Differentially private conformal prediction method for uncertainty quantification with statistical efficiency under privacy constraints.
World-Value-Action model for vision-language-action agents that enables implicit planning and long-horizon reasoning beyond direct action prediction.
OpenAI expands Codex to enterprises through partnerships with GSIs, reaching 4M weekly developers. Codex deployment across software development lifecycle workflows.
Full-stack Python framework compiling Python to JavaScript for browser execution with unified state management.
Plugin for Claude Code that auto-generates session titles and colors to distinguish parallel LLM sessions.
Research on safety-filtered LLM pretrains showing probabilistic censoring gaps across models from multiple labs.
Managed vector index service with dynamic hybrid search combining dense embeddings and BM25, achieving 12× compression and 10× faster queries with per-query fusion.
Open-source agentic framework for quantitative finance with self-improving portfolio strategies and backtesting.
Study examining how AI tools like Cursor enable improved workflow and focus for neurodivergent developers.
DevArch 2.0 provides directives and agents for Claude Code enforcing engineering discipline with automated quality gates and testing.
MCPfinder aggregates and helps discover/install Model Context Protocol servers from multiple registries for agent configuration.
1Password shares experience using AI agents to refactor a multi-million-line Go monolith, covering successes and failures.
Opinion on token inflation in LLM pricing and feature bloat driven by AI capabilities.
Astrolabe declarative macOS configuration framework inspired by SwiftUI for MDM state convergence.
Multi-provider LLM systems become harder to operate when adding reasoning capabilities due to infrastructure and abstraction challenges.
Verus tool for static verification of Rust code using theorem proving without runtime overhead.
Cursor CLI agent adds Debug Mode for hypothesis-driven bug fixing and /btw support for side questions.
Open-source geospatial vector/raster server using Python, PostGIS, and OGC APIs. Built with AI assistance.
Openheim is an open-source LLM agent written in Rust that operates as CLI, REPL, or HTTP server.
Web tool using Gemini 2.5 Pro to simplify and explain legislation and executive orders.
Developer compares open-source AgentHandover project with OpenAI's Chronicles feature for agent screen monitoring.
LeWorldModel: stable end-to-end JEPA for learning world models from raw pixels using minimal loss terms. Original ML research on self-supervised learning.
2-layer transformer model implemented in 6502 assembly running on Commodore 64. Demonstrates transformer architecture feasibility with ~25K parameters.
Comrade: open-source AI workspace for security teams with transparency and extensibility focus. GitHub project for agent-based workflows.
Prism v11.0 cognitive architecture for AI agents with memory and causal reasoning. Product claims unclear with limited technical details.
Discussion of embedding AI agents directly in software applications rather than standalone systems, with practical examples from Feldera's infrastructure use.
ARGOS is an open-source autonomous AI infrastructure agent for monitoring and managing server fleets via natural language, built for self-hosted use without cloud dependency.