Breakthrough in AI Solving Math Conjectures: Peking Univ. Team's Exploration
Peking University AI framework autonomously discovered and formally verified solution to open problem in commutative algebra using 19,000 lines of Lean 4 code.
Peking University AI framework autonomously discovered and formally verified solution to open problem in commutative algebra using 19,000 lines of Lean 4 code.
Multi-model workflow orchestrator with YAML-defined chains, parallel execution, visual canvas builder, and MCP/REST/CLI interfaces supporting tool-based agents.
Safetensors joins PyTorch Foundation as hosted project to secure model distribution and prevent arbitrary code execution in agentic solutions.
Local prompt injection detection tool for AI agents using MCP and function calling, runs ONNX model offline without API calls.
Announcement of Claude Mythos Preview availability on Google Cloud Vertex AI for select customers as part of Project Glasswing.
Browser-based private AI workspace running models via WebGPU with zero server communication, local storage, and no API key requirements.
PostgreSQL MCP server providing schema awareness to AI agents and coding assistants from offline snapshots without production database credentials.
Research note on Mamba-3, a state space model architecture improving on transformers' quadratic attention costs for efficient LLM inference in agentic applications.
A2CN: open protocol for safe agent-to-agent commercial negotiation enabling procurement agents to transact deals.
COBOL-based AI agent chatbot with tool use and agentic loop, integrating with modern LLMs via OpenRouter.
Reading notes on signature method for feature engineering from sequential event data in machine learning.
Self-hosted AI assistant with voice, vision, RAG, and web search running entirely on-device without cloud.
Open-source, self-hosted form backend alternative to SaaS solutions like Formspree with email/webhook delivery.
AI video generation tool converting text and images to videos for content creation platforms.
Automated decensoring tool reduced Google's Gemma 4 refusal rate from 98% to 47% in 24 minutes on laptop.
Open-source infrastructure for building AI agents and internal software with managed database, auth, and RBAC.
SharpSkill is an interview preparation platform for coding assessments covering React, Node.js, and other tech stacks.
Analysis of Entire.io startup and its Checkpoints open-source CLI tool providing observability layer for AI coding workflows.
Investigation into false benchmark claims in MemPalace, an open-source AI memory project. Exposes fabricated scores and questions attribution to actress Milla Jovovich.
Python real-time engine with sub-1ms jitter for industrial control, auto-generates REST APIs and MCP for LLM agent integration.
Anthropic provides Mythos model to major tech companies for cybersecurity testing and vulnerability discovery.
Open-source framework for AI SRE agents that integrate 40+ infrastructure tools to autonomously investigate and resolve production incidents.
Video presentation on GitOps relevance and practices in systems managed by AI agents, from FluxCon conference.
Yu is a sandboxing tool that isolates Claude Code and Codex execution to prevent credential exposure from compromised code or dependencies.
GLM-5.1 is a 754B parameter open-source LLM that demonstrates improved reasoning and multi-modal capabilities like unprompted SVG+CSS generation.
Analysis of cognitive load and limitations when managing multiple parallel AI agents, focusing on human-in-the-loop costs beyond throughput metrics.
Enterprise authorization system for Model Context Protocol (MCP) servers using centralized identity providers. Addresses deployment challenges in large organizations.
Research on improving code reviews by adding semantic analysis layer to local LLMs, providing contextual function/type information beyond diffs.
Google releases offline-first dictation app using Gemma-based ASR models. Open-source LLM application for speech recognition on consumer hardware.
Tool that detects blind spots in AI coding agent pull request reviews by analyzing API and database boundary changes. Addresses integration testing gaps.
Voice-first AI planning tool with MCP integration. Conversational AI agent for strategic planning that generates structured documents in real time.
GPU-resident vector database (~300KB executable) supporting 12M vectors with ~10ms query latency, TCP interface, no dependencies.
Security analysis: 25,000+ publicly exposed Ollama instances found in April 2026, 22x increase from September 2025, raising infrastructure security concerns.
Open-source lakehouse demo using DuckDB, dlt, and dbt. Complete runnable example of ELT pipeline with parquet files and analytics transformation.
KOS Protocol: open standard for publishing machine-readable verified facts with provenance tracking and freshness decay. Addresses AI hallucination via structured data.
Research on fine-tuning LLMs for epistemic reasoning using Navya-Nyaya logic. Addresses hallucination and brittleness in LLM reasoning capabilities.
ReVEL hybrid framework uses LLM-guided iterative evolution with structured performance feedback to design effective heuristics for NP-hard problems.
Framework identifies algebraic structures in combinatorial optimization problems, constructs quotient spaces to reduce search space and improve solution quality.
PaperOrchestra multi-agent framework automates AI research paper writing by transforming unstructured materials into submission-ready LaTeX manuscripts.
MMORF multi-agent framework uses language models with specialized agents for multi-objective retrosynthesis planning balancing quality, safety, and cost.
MedGemma 1.5 4B model expands medical capabilities with high-dimensional imaging (CT/MRI/histopathology), anatomical localization, and improved document understanding.
LLM-based sequential clinical diagnosis system models uncertainty-guided evidence acquisition over time using diagnostic trajectory learning.
Kolmogorov-Arnold Fuzzy Cognitive Maps extend neuro-symbolic modeling to handle non-monotonic causal dependencies in complex dynamic systems.
IntentScore is a plan-aware reward model trained on 398K offline GUI interactions to evaluate and score actions for computer-use agents across multiple operating systems.
Instruction-tuned LLMs parse and mine unstructured HPC system logs from heterogeneous sources to extract patterns and diagnose operational issues.
ClawsBench benchmark evaluates LLM agents on realistic productivity tasks (email, scheduling, documents) in simulated multi-service environments with stateful workflows.
AttriBench: Demographically-balanced benchmark for measuring attribution bias in LLMs when attributing quotes to original authors.
Framework for translating governance norms into enforceable runtime guardrails for agentic AI systems with multi-step execution.
Evolutionary theory simulation of how alignment affects populations of AI models over time and belief propagation dynamics.
Reward decomposition approach to disentangle pressure capitulation from evidence blindness in LLM sycophancy behavior.