Show HN: AgentWeb – Free business directory API for AI agents (11M+ businesses)
Free business directory API with 11M+ businesses, full-text search and geo-location for AI agents to query real-world business data reliably.
Free business directory API with 11M+ businesses, full-text search and geo-location for AI agents to query real-world business data reliably.
Framework for governing autonomous AI agents in production environments, addressing safety and control.
Research on LLM-generated passwords being predictable, exposing security vulnerabilities in password generation tasks.
Discussion about how LLM prevalence affects open source contributor motivation and reputation, questioning whether AI-generated code diminishes accomplishment recognition.
Collection of 380+ agent skills from Anthropic, Google, Vercel, Stripe and other teams for use with Claude, Codex, Gemini CLI and similar platforms. Real-world skills, not auto-generated.
MIT study reports agentic AI moving mainstream with capabilities outpacing governance; notes OpenAI hired OpenClaw framework creator.
Analysis of legal and regulatory risks from AI agents equipped with crypto wallets for autonomous transactions and hiring.
Anthropic offers 6 months free Claude Max 20x to open-source maintainers and contributors through rolling application review.
Kill switch governance tool for autonomous AI agents using LangChain, preventing agent sprawl and cost overruns from runaway recursion.
Discusses designing LLM applications as systems built around models, planning for quarterly model swaps rather than model-centric architecture.
Claim that AI agent supply chain has problems with proposed solutions, no substantive details provided.
Case study of AI voice agents deployed in hotels analyzing 15,910 real guest interactions and lessons learned from deployment.
Personal experience using AI-assisted coding with GLM 4.7/5 models and Claude CLI for debugging and development tasks, particularly Vulkan code.
Open source static documentation generator from single Markdown file with AI agent support via llms.txt format. Includes CLI tool, dark mode, search indexing.
OpenAI and Amazon announce multi-year strategic partnership with $50B investment. Joint development of Stateful Runtime Environment for enterprise AI.
Amazon Bedrock launches Stateful Runtime Environment for AI agents, powered by OpenAI models. Enables multi-step reliable agent execution with tool integration and operational controls.
Theoretical study of Takeuchi's information criterion as a generalization measure for deep neural networks in the neural tangent kernel regime.
Offline reinforcement learning approach using physics-informed constraints via PDEs to improve value function estimation in goal-conditioned policies.
ParamMem augments language agents with parametric reflective memory to improve reasoning by increasing reflection diversity.
Differentiable approximation to zero-one loss via hypersimplex projections for end-to-end optimizable classification models.
Memory-efficient optimizer suite (FlashOptim) for training large language models with reduced accelerator memory requirements.
Semi-supervised alignment of vision and language models using optimal transport with minimal paired samples.
Dataset compression approach enabling efficient distribution of training data to clients with limited hardware resources.
Survey of neural routing solvers using deep learning to replace handcrafted heuristics for vehicle routing optimization problems.
Large-scale text-to-SQL dataset (SQaLe) with 135k+ database schemas for training generalizable natural language to SQL conversion models.
Query-aware chunk compression method for RAG systems that dynamically optimizes document retrieval and generation based on input queries.
Multimodal fusion of biological LLMs with state space models for improved RNA interaction prediction.
Analysis of backdoor vulnerabilities in multimodal diffusion language models with self-purification defense mechanism.
Benchmark evaluating LLMs on financial knowledge through exam questions and practical business reasoning scenarios.
Training-free framework for multimodal information retrieval using MLLMs without large datasets or pre-training fine-tuning.
Large-scale hypothesis screening of biological foundation models to understand what geometric and topological structures they learn from gene expression data.
Multimodal LLM framework analyzing video ad effectiveness by examining the first three seconds using visual, audio, and text analysis.
Analysis of function vectors in LLMs showing they lack invariance across input formats despite targeting same concepts.
Testing framework to diagnose whether MLLMs genuinely read text in images or rely on parametric shortcuts in prompts.
GetBatch: Object store API elevating batch retrieval to first-class operation for efficient ML data loading across storage clusters.
veScale-FSDP: Enhanced FSDP distributed training framework supporting block-wise quantization and non-element-wise optimizers.
Empirical study of latent reasoning methods under weak and strong supervision for multi-step LLM reasoning.
Uncertainty-aware policy steering using VLM verifiers to adapt robot behaviors by selecting aligned action samples.
VeRO: Evaluation framework for assessing coding agents that optimize other agents through iterative edit-execute-evaluate cycles.
Flow matching generative model that adapts to manifold structures, offering simulation-free alternative to diffusion models.
Theoretical analysis of scaling limits from shallow Bayesian neural networks to Gaussian processes with scalable inference methods.
Dynamic dense retrieval with routing strategy for adapting information retrieval models across domains without full retraining.
CourtGuard: Model-agnostic multi-agent framework for zero-shot LLM safety policy adaptation using retrieval-augmented debate.
Search-P1: Path-centric reward shaping for training agentic RAG systems with improved sample efficiency via RL.
Item Response Theory approach to correct systematic rater biases in human evaluations for AI model assessment.
SideQuest: Model-driven KV cache management technique for long-context agentic reasoning tasks with multi-hop retrieval.
Hybrid ML framework combining autoencoders and transformers for accelerator beam diagnostics simulations.
Trie-based constrained decoding optimization for LLM generative retrieval on accelerators. Improves business logic constraints in recommendations.
dLLM: Unified framework for diffusion language models. Standardizes components across research implementations for reproducibility.
GR4AD: Production generative recommendation system for large-scale advertising using LLMs. Architecture, learning, and serving optimization.