Show HN: NeoMud – A multiplayer dungeon game with AI agents that QA and playtest
MUD game built with Claude Code featuring agent pipeline: game-designer agent proposes balance changes, playtest agent QA-tests gameplay.
MUD game built with Claude Code featuring agent pipeline: game-designer agent proposes balance changes, playtest agent QA-tests gameplay.
Discussion questioning whether subagents in AI coding assistants create manager-like behavior, reducing flagship models' reasoning autonomy.
LLM-based email parser tool handling brittle format variations with schema validation, retries, and backend integration.
Conversational AI agent for Kubernetes cluster operations. Combines visual dashboard with agentic workflows to reduce context switching between monitoring and CLI tasks.
Open-source maintainers face harassment from AI-generated code contributions. matplotlib and other projects implement human review policies for AI-submitted code.
iPad agentic coding tool using Claude that autonomously reads codebases, plans changes, edits files, and pushes to GitHub. 7 integrated tools execute locally with token-streaming API.
Hopsworks platform as runtime for coding agents. Analysis of container vs wrapper approaches and MCP integration for Claude Code, Codex, Gemini.
Voice-first AI language learning application built with Next.js and Gemini APIs. Technical discussion of browser-native lip sync implementation challenges.
Tool that generates unified context files (CLAUDE.md, .cursorrules, etc.) for AI coding tools from a single codebase scan, solving maintenance overhead.
Platform for building and deploying autonomous AI agents across decentralized networks. Combines learning environment with infrastructure for cross-chain finance and trusted execution.
Safety guardrails tool intercepting dangerous commands from both humans and AI agents. Works universally across shells with blast radius detection and context analysis.
Side project: 1v1 strategy game where LLMs compete. Results show LLMs can generate functioning bots but underperform vs. human players.
Article exploring optimal human-agent collaboration in software development. Discusses whether developers should inspect all code or focus on outcomes.
Open-source helpdesk software version 7.0 adds AI features without vendor lock-in. Supports pluggable LLMs while maintaining data sovereignty.
Three-stage curriculum learning framework for distilling chain-of-thought reasoning from large LLMs into compact student models while preserving interpretability.
cc-Shapley method for measuring multivariate feature importance in ML models using causal context.
Zatom-1 open-source foundation model for unified generative and predictive learning on 3D molecules and materials.
Reference-guided fine-tuning approach for reinforcement learning on mathematical reasoning tasks with sparse rewards.
AOI framework for training LLM agents to diagnose cloud infrastructure failures using failed trajectories as learning signals, addressing safety and data constraints.
Neural solver for vehicle routing problems using distance representation learning for asymmetric real-world scenarios.
Analysis of why linear RNNs achieve transformer-level parallelizability compared to nonlinear RNNs for language modeling.
Learning approach for Dubins Traveling Salesman with neighborhoods using knowledge distillation from expert trajectories.
Theoretical research on game theory attractors and replicator dynamics, characterizing learning equilibria.
Research paper on safety mirage in vision language models: spurious correlations in safety fine-tuning and mitigation via machine unlearning.
Evaluates LLM fault localization capabilities on code changes, assessing semantic program reasoning beyond syntax and lexical features.
Benchmark evaluating LLMs on rigorous causal inference tasks involving statistical pitfalls relevant to medicine, economics, and public policy.
ShIOEnv: Gymnasium environment for grammar-constrained synthesis and modeling of command-line interface behavior with shell input-output data.
Novel framework using multi-kernel Boolean parameters for weight binarization in LLMs to improve efficiency without full-precision latent weights.
SealQA benchmark for evaluating search-augmented LLMs on fact-seeking questions with conflicting or noisy web search results.
EDINET-Bench: New benchmark dataset evaluating LLMs on complex financial document analysis tasks using Japanese financial statements.
MuRating framework transfers English data-quality signals to score documents in 17 languages for multilingual LLM pretraining.
Research generalizes EDM to arbitrary-noise diffusion models, analyzing design space beyond Gaussian noise for image restoration tasks.
Quantum EM algorithm for training quantum Boltzmann machines that circumvents barren plateau problem in quantum machine learning.
Study shows LLM ranking systems are sensitive to small changes in preference data, with top model rankings changeable by dropping few preferences.
Comprehensive evaluation of how weight and activation quantization affects model bias across stereotypes, fairness, toxicity, and sentiment.
Research on optimal alignment of acoustic and linguistic representations in pre-trained models for automatic speech recognition.
BabyHuBERT self-supervised speech model trained on 13,000 hours multilingual child-centered recordings for speaker segmentation.
Framework combining diffusion models with impedance control for robot learning in contact-rich manipulation tasks.
Noise-to-Notes reformulates automatic drum transcription as conditional generative task using diffusion modeling.
BridgeDrive applies diffusion-based planning with expert behavior anchors for closed-loop autonomous driving trajectory planning.
BeyondBench framework uses algorithmic problem generation for contamination-resistant evaluation of reasoning in language models.
SphereAR addresses variance collapse in continuous-token autoregressive image generation by constraining latents to hypersphere.
Theoretical analysis of quantitative convergence of shallow neural networks trained via gradient descent to Gaussian processes.
Research on NVFP4 quantization approach for efficient LLM pretraining, reducing compute and energy requirements for frontier models.
VidGuard-R1 uses reasoning MLLMs and reinforcement learning to detect AI-generated videos with human-interpretable explanations.
Research on self-supervised novel view synthesis identifies transferability as key criterion for true NVS capability across video sequences.
Safety filtering for reinforcement learning using Control Barrier Functions to enforce dynamic safety constraints during training.
Multimodal foundation model for accelerating numerical simulation of stochastic differential equations via neural network-based error correction.
CoRPO: Adds correctness bias to GRPO reinforcement learning for improved reasoning and generalization in LLMs.
ObAct: Imitation learning framework for active vision in dual-arm robots using 3D Gaussian Splatting.