Efficiency of Parallel and Restart Exploration Strategies in Model Free Stochastic Simulations
Analysis of parallelization and restart mechanisms for exploration in model-free reinforcement learning and rare event estimation.
Analysis of parallelization and restart mechanisms for exploration in model-free reinforcement learning and rare event estimation.
Transformer-based routing framework for single-agent NP-hard combinatorial optimization in dynamic IoT network environments.
Survey of LLM-based human-agent collaboration systems addressing reliability, complexity, safety and trustworthiness challenges in autonomous agents.
Multi-armed bandit algorithms for optimal policy learning under network interference in sequential intervention assignment.
Re-evaluates instruction-guided navigation systems, analyzing geometry vs LLM contributions with training-free variants for robot navigation.
Scaling continuous-time consistency models for fast diffusion in large-scale text-to-image and video generation tasks.
Optimization framework for performing downstream tasks on learned Riemannian data manifolds with low-dimensional structure.
Framework exposing 'Logic Inertia' problem in LLMs where models fail on structural perturbations of rule-based systems.
Algorithm for computing token likelihood ratios between language models with different tokenizers for knowledge distillation.
Study on structured width pruning of Llama models showing trade-offs between knowledge retention and instruction-following capability.
Method for unlearning unsafe concepts in image generation models through prompt embedding redirection.
Research on evasion attacks against LLM-based code vulnerability detectors using syntax-preserving transformations on C/C++ benchmarks.
Contrastive learning framework (CLAMP) for pretraining 3D multi-view representations in robotic manipulation policies via action-conditioned self-supervision.
Systematic study of cross-objective interference in multi-objective LLM alignment, showing performance improvements on some objectives cause degradation on others.
Statistical method to detect significant model degradations in LLMs after optimization techniques like quantization, with robustness analysis at zero temperature.
Particle filtering algorithm for robotic state estimation trained with single-step objectives rather than end-to-end sequence modeling.
Survey of meta-learning and meta-reinforcement learning techniques enabling rapid adaptation to new tasks with minimal data, tracing path to adaptive agents.
LABBench2 benchmark for evaluating AI systems performing biology research, covering foundation models, agentic hypothesis generation, and autonomous labs.
Multi-agent reinforcement learning framework for wind farm flow control balancing power output with structural load constraints using Independent Soft Actor-Critic.
Comprehensive survey of security threats and defenses in LLM-based AI agents, organizing attacks by architectural layers including memory, tools, and multi-agent interactions.
Proposes Path-Lock Expert architecture to cleanly separate reasoning and non-reasoning modes in hybrid-thinking LLMs by isolating parameters, reducing reasoning leakage in inference.
Empirical study showing simple system-prompt self-orchestration outperforms agent orchestration frameworks for procedural LLM tasks.
Study of perceptual bandwidth bottleneck in VLMs proposing active visual reasoning via sequential experimental design for fine-grained reasoning.
Anon optimizer improves adaptation across diverse landscapes by decoupling pre-conditioner adaptivity, generalizing better than Adam on various architectures.
ANO optimizer addresses PPO's hard clipping limitation with principled design space for robust policy optimization in RL and LLM alignment.
CreativityBench benchmark evaluates LLM creative reasoning through affordance-based tool repurposing and non-canonical object usage.
RLDX-1: Vision-Language-Action robotic policy model addressing limitations in motion awareness, memory, and physical sensing for complex tasks.
Technical explanation of how LLMs work: tokenization, embeddings, transformer layers, self-attention, token prediction. Educational deep dive with code examples.
0xBitNet enables running BitNet b1.58 ternary LLMs via WebGPU in browsers and native apps using custom WGSL kernels with TypeScript, Rust, and Python bindings.
SDK enabling AI agents to pay for API calls using HTTP 402 protocol. Payment infrastructure allowing agents to autonomously purchase services.
Job search platform using AI to improve hiring experience beyond mass applications. LLM-based recruitment tool with Google Embeddings integration.
Hunk: terminal diff viewer for agent-authored changesets built on OpenTUI with review-first workflow for agentic coders, supports agent-context and skill paths.
sqlalchemy-redshift Python library revived for Amazon Redshift SQLAlchemy dialect with PyPI distribution supporting redshift_connector and psycopg2.
Simplex uses ChatGPT and Codex to accelerate software development, reporting 70% faster screen development and 40% faster design iterations.
Trump administration considering voluntary frontier AI model vetting system using existing CAISI/CISA tools, enabling labs to provide government early access to models with cyber capabilities.
Research on on-policy LLM distillation (2025): training smaller domain-expert models outperforming larger generalist models through stacked training stages including perception, retrieval, planning, and execution.
gpu.fund tracks GPU cloud rental prices across providers (H100s, 4090s, 3090s) to help users find cheapest available hardware for training and inference.
Brydg is an AI hiring platform that automates interview scheduling, resume ranking, and offer management via Google Meet.
Anthropic doubles Claude Code usage limits following a compute deal with SpaceX for increased developer access.
Security researcher reports Google Chrome silently downloads 4GB on-device AI model without user consent.
Study shows students using LLMs for programming assignments experience reduced learning and understanding despite perceived improvement.
Braintrust, an AI evaluation startup, confirms AWS breach exposing customer API keys for cloud AI model access.
Private self-hosted AI assistant running locally on user machines via Ollama, integrating with messaging platforms without cloud data transmission.
Terminal coding agent for DeepSeek V4 that streams reasoning blocks, edits local code with approval gates, and auto-selects models.
Production-ready open-source framework for sovereign AI agents with cryptographic identity, persistent memory, and constitutional governance.
Grok introduces Connectors, integrations enabling end-to-end automation with email, slides, calendar, and spreadsheets without manual copy-pasting.
Open-source self-hostable background agent platform inspired by Ramp Inspect, enabling AI agents to work autonomously on GitHub repos.
Aion is a collaborative coding game where players and AI agents propose diffs compiled live every 5 minutes via MCP-enabled agents.
Genosyn is an open-source self-hostable platform for running autonomous AI companies with AI employees having constitutions, skills, and scheduled routines.
Pay.sh is a payment layer for HTTP APIs handling x402 stablecoin payment challenges, enabling autonomous CLI tool access to gated endpoints.