Docker Sandboxes and Docker Agent
Docker Sandboxes enables AI agents to autonomously handle multi-disciplinary development tasks. Frames agents replacing context-switching across product/design/engineering roles.
Docker Sandboxes enables AI agents to autonomously handle multi-disciplinary development tasks. Frames agents replacing context-switching across product/design/engineering roles.
API enabling AI agents to handle document signing workflows end-to-end. Solves agent workflow bottleneck with markdown-to-PDF and URL-based PDF signing.
AI-powered landing page generator producing copy, layout, and design automatically. LLM application with marketing focus, limited technical novelty.
Comparative analysis of LLMs for code generation and debugging, examining performance across reasoning, code generation, and general understanding tasks.
Video benchmarking LLMs on Eleusis game of science task. Evaluates LLM reasoning capabilities.
CLI tool enabling AI agents to control web browsers using existing login sessions across 36 platforms without APIs or scrapers.
AllocDB is a deterministic resource-allocation database built with Codex using strict architectural principles, tested with Jepsen and KubeVirt infrastructure.
Configuration system for Cursor AI editor defining custom rules and behaviors for code generation via .cursorrules files.
Neural network-based CPU implementation running on GPU with differentiable computation graph. Conceptual project exploring gradient descent optimization of programs.
Self-hosted visualization tool using AI agents with GitHub Copilot CLI to generate and organize dashboards from Jira data. Open source developer tool.
Announcement of GPT-5.3-Codex-Spark model for real-time coding in Cursor IDE, 1000+ tokens/sec, 128k context window, text-only. Details sparse, appears promotional.
Shard automatically decomposes complex coding tasks into parallel DAG sub-tasks, allowing multiple AI agents to work simultaneously with zero merge conflicts.
News aggregation site converting AI security research papers into articles, covering LLM deception risks, agent architectures, and attack surface mapping.
AI tools lower barriers to open source contributions by helping developers understand codebases and projects, shifting focus from syntax mastery to problem intent.
MCP server for managing Meta's Threads from Claude, built with Claude Code. Enables social media automation through AI agent integration.
Neuroscope tool providing real-time interpretability into LLM internal representations. Developer tool for understanding LLM behavior.
Opinion piece on how LLMs enable overconfident employees to obscure lack of competence. Commentary on LLM societal impact.
Critical analysis comparing LLMs to epicycles in astronomy, questioning whether intelligence is the appropriate metric for evaluating current language models.
Port42: SwiftUI app enabling AI companions to build interactive UIs and act on macOS. Open source developer tool with live code demo.
Agent harness concept: software infrastructure wrapping LLMs/agents for orchestrating tools, memory, workflows. Technical introduction with architectural focus.
Agentic Trust Framework: open security specification for Zero Trust governance of autonomous AI agents. Standards and governance for agent deployment.
BotStadium: research platform simulating AI agent behavior through competitive sports predictions. Agent behavior analysis and testing platform.
LLM Architecture Gallery: curated collection of architecture diagrams and specifications for major LLMs. Technical reference resource.
Multi-VLM ensemble method using vision and language modalities to select complementary models for efficient visual reasoning.
Composite attack on LLM safety alignment where multiple LoRA adapters appear benign individually but suppress safety when composed.
Defense mechanism against adversarial patches in Vision Transformers using token segregation and randomized transformations.
Hierarchical LLM-based approach for fine-grained multi-table retrieval using compositional reasoning instead of coarse-grained similarity matching.
In-context learning strategy for CAD code generation using design-specification tiling to improve LLM performance on domain-specific tasks.
Foundation model-guided approach for virtual immunohistochemistry staining from H&E images to accelerate pathology diagnostics.
Multimodal recommendation framework using anchor-based alignment in projection space to prevent modality collapse and ID dominance.
Novel 3D molecule generation framework using vector-field representations to address modality entanglement and geometry-chemistry constraints.
Self-supervised system for robots to detect and recognize novel objects from human video demonstrations without prompt engineering.
Efficient optimization technique addressing long-tail distribution problem in LLM-based sequential recommender systems.
Multimodal foundation model for Earth observation using temporal training objectives robust to variable-length satellite and sensor data.
Method for injecting auxiliary visual features into vision-language-action models to improve geometric understanding and temporal reasoning for robotic manipulation.
Model distillation approach compressing 2B vision-language retriever into 70M text-only encoder for efficient document retrieval.
AI-based framework transforming global weather forecasts into fine-grained wind field predictions and infrastructure failure probabilities for tropical cyclones.
Research on using vision-language models for detecting and localizing forged images, studying how VLM priors affect forgery detection performance.
Data-efficient MRI reconstruction strategy using diffusion probabilistic models with pre-training and fine-tuning.
Adaptable fraud detection system handling adversarial attacks in resource-constrained environments with multiple risk modules.
Framework (ARL-Tangram) optimizing resource efficiency in agentic RL by dynamically allocating external compute resources.
Surgical world model using controllable video generation for simulating surgical actions with precise tool-tissue control.
Dataset and baseline for real-time screw classification in industrial automation and robotic systems.
Diagnostic benchmark (ESPIRE) for evaluating vision-language models on embodied spatial reasoning tasks.
GNN approach for precoder learning in cell-free wireless systems accounting for dynamic user-access point associations.
Empirical study of federated few-shot learning on neuromorphic hardware using spike-timing-dependent plasticity.
Convergence analysis of functional learning methods for contextual stochastic optimization problems.
Research on interpretable multimodal concept bottleneck models ensuring faithful explanations through proper concept detection.
Provable multi-agent reinforcement learning in partially observable stochastic games leveraging information sharing among agents.
Introduces diffusion models as expressive variational posteriors for black-box inference in latent variable models.