Vision Banana: Image Generators Are Generalist Vision Learners
Vision Banana: Multi-task vision model achieving state-of-the-art zero-shot transfer across 2D/3D vision tasks using generative approach.
Vision Banana: Multi-task vision model achieving state-of-the-art zero-shot transfer across 2D/3D vision tasks using generative approach.
agentcall.dev: Platform enabling Claude Code and other agents to join video calls as participants with voice, screen-sharing, and live coding.
Anthropic's Mythos enterprise security AI model reportedly breached by unauthorized users through third-party vendor access. Model was not intended for public release.
Elon Musk's xAI exploring partnership with Mistral and Cursor to compete with Anthropic and OpenAI. SpaceX has option to acquire Cursor.
Google announced Chrome as agentic workplace platform with Auto Browse for autonomous task completion, Chrome Skills for AI workflows, Gemini integration, and enterprise controls via Chrome Enterprise Premium.
agents-cli tool for building AI agents on Google Cloud platform with coding agent integration.
Computational framework for fast prime number generation using spectral analysis and signal processing on MacBook Pro.
JibarOS is Android 16 fork adding shared inference runtime for on-device AI with system service, native daemon, pluggable backends, and Binder AIDL for text/audio/vision capabilities.
Research on LLM benchmark ranking inconsistency; proposes train-before-test method improving agreement across 61 models and 24 benchmarks.
PiClaw v1.8.5 release notes describing breaking change in pi-coding-agent tool parameter semantics affecting extension-registered tools.
Cross-platform desktop widget displaying Claude Code API rate limits and token usage in real-time with historical analytics.
Corral visualization tool for measuring how LLM-based AI scientists reason through epistemological graphs of research traces.
Empirical study testing 40 Claude prompt codes to determine which actually change reasoning vs. output style; 7 showed meaningful differences.
Opinion piece arguing AI agents need inter-agent communication and signaling capabilities beyond single-task, single-user designs.
Open-source IDE for managing multiple AI coding agent sessions with parallel execution, cost tracking, and multi-provider support (Claude, OpenAI, Gemini).
NVIDIA's Asset Harvester converts driving logs into 3D assets for autonomous vehicle simulation using neural scene reconstruction.
AISBF open-source framework providing unified proxy, intelligent routing, and failover across multiple AI providers (Google, OpenAI, Anthropic, Ollama).
Developer used Claude Code to build an AI system that understands and executes 6502 assembly language with real-time feedback loops.
DHS researchers demonstrated jailbreak techniques against popular LLM safeguards to lawmakers, highlighting vulnerabilities in AI systems.
Study showing LLM health advice contains significant errors and fabricated citations despite appearing authoritative.
Control plane for autonomous AI agents that prevents unbounded retry loops through budgets, policy checks, rollback evidence, and audit trails.
Research analyzing tool overuse phenomenon in LLMs, examining why models unnecessarily invoke external tools over internal knowledge.
Framework for governance and evaluation of generative AI in learning-intensive domains addressing proxy failure problems.
Research proposing feature-free algorithm selection using pretrained text embeddings without domain knowledge.
Research on data augmentation and resampling strategies for transformer models addressing class imbalance in educational assessment.
Research on explainable LLM-based triage for anti-money laundering with evidence retrieval and counterfactual verification.
ThermoQA benchmark with 293 engineering thermodynamics problems in 3 tiers evaluating reasoning in 6 frontier LLMs against programmatic ground truth.
EvoForest machine learning paradigm uses open-ended evolution of computational graphs to discover transformations and structures rather than just optimize parameters.
Framework for interpreting temporal concept evolution in LLM agents through conformal methods to understand sequential decision-making mechanisms.
Position paper on integrating learning theories into explainable AI (XAI) for large complex models to improve transparency.
Autonomous LLM agent framework for materials science theory development that generates equations, writes code, and validates theories without intervention.
Study identifying reliability risks in LLMs across different numerical precision formats (bfloat16, int8) causing output disagreements.
OpenCLAW-P2P v6.0 decentralized platform where autonomous AI agents publish, peer-review, and improve research papers without human gatekeepers.
SkillGraph uses execution-transition graphs mined from 49k workflows to recommend tool sequences for LLM agents, outperforming semantic-only methods.
Prism unifies memory substrates for multi-agent AI systems combining vector memory, graph relations, file persistence, and evolutionary search for open-ended discovery.
AI Telco Engineer framework uses agentic LLMs to autonomously generate, evaluate, and refine wireless communication algorithms iteratively.
MIRROR benchmark evaluates metacognitive calibration in 16 LLMs across 250k instances across 8 experiments measuring self-knowledge and decision-making.
DrugKLM combines knowledge graphs with LLM reasoning for drug repurposing and therapeutic candidate prioritization with mechanistic grounding.
Emergence Transformer explores temporal attention mechanisms in transformers and their role in emergent phenomena, with theoretical analysis of long-range interactions.
JTPRO framework optimizes LLM agents with tools by jointly refining prompts and tool schemas to reduce mis-selection and improve slot instantiation in large tool libraries.
Forage V2 extension enabling autonomous agent organizations to evolve knowledge and transfer learnings across multiple expeditions in open-world tasks.
Active inference-based computational model for understanding how autonomous vehicles and road users resolve space-sharing conflicts through uncertainty reduction.
Framework addressing AI presumptuousness in legal applications, teaching systems when not to decide based on insufficient evidence in unemployment insurance adjudication.
Framework for iterative creative game generation using LLMs with mechanic-aware optimization, addressing brittle runtime behavior and weak experience accumulation.
Diagnostic framework for evaluating AI-generated peer reviews at concern level rather than verdict level, assessing alignment with review rationale.
Framework enabling AI agents to perform causal reasoning through experimentation and hypothesis-space restructuring using architectural scaffolding.
AI system for hospital quality improvement factor discovery, automating identification of modifiable contributing factors in healthcare optimization.
EvoAgent framework for LLM agents with structured skill learning and hierarchical sub-agent delegation, enabling continuous skill generation through user feedback.
Hierarchical Preference Optimization method extending DPO for complex reasoning tasks in LLMs, providing granular feedback on multi-step solutions.
Architecture for enterprise AI agents in regulated domains using stateless decision memory instead of retrieval-augmented pipelines, enabling deterministic replay and auditability.