Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
Dynamic scheduling system for efficient large model training across GPU clusters. Addresses training efficiency and resource utilization.
Dynamic scheduling system for efficient large model training across GPU clusters. Addresses training efficiency and resource utilization.
Self-trained fine-tuning paradigm for LLMs on table understanding tasks like NL-to-Code and data cleaning. Reduces need for expensive human labeling.
Continual learning technique combining parameter-efficient fine-tuning with vision transformers to prevent catastrophic forgetting. Addresses sequential task adaptation.
Method for training LLMs to explain their own internal activations using natural language probes. Advances LLM interpretability research.
Transformer attention mechanism compressed to run in under 2MB memory for IoT and wearable devices. Enables NLP deployment on ultra-constrained hardware.
Study redefining non-IID data heterogeneity in federated learning by migrating from label to embedding-level task-specific distributions.
Learning dynamically-inspired bases for Koopman and transfer operator approximation in complex nonlinear dynamical systems.
CounterLogic benchmark evaluating LLM reasoning in counterfactual scenarios where context contradicts parametric knowledge.
Method for text-to-image diffusion models to handle contextually contradictory prompts where concepts implicitly negate each other.
Large-scale benchmark with 1,507 real-world vulnerabilities evaluating AI agents' dynamic cybersecurity capabilities at scale.
Masked conditional generative model for peptide discovery that predicts aggregate morphology for biomedical material design.
BuilderBench benchmark for evaluating intelligent agents' ability to learn through interaction and exploration beyond training data.
Reinforcement learning approach using information gain-based rewards to optimize LLM agents for multi-turn search with tool use.
Application of diffusion models to semantic communications in 6G wireless systems for meaning-centric data transmission.
Multimodal model addressing modality imbalance and noise in e-commerce product understanding with dynamic balancing.
Framework for incorporating inference delays into diffusion policy learning for robotic control in dynamic environments.
Analysis of positional encoding impact on Transformer generalization and robustness in in-context regression, showing PE enlarges generalization gap.
ASK framework addresses gradient locality bottleneck in audio-text retrieval by incorporating external knowledge injection in dual-encoder architectures.
1S-DAug introduces one-shot generative data augmentation synthesizing diverse image variants from single examples for improved few-shot learning generalization.
KDFlow is a knowledge distillation framework for compressing large language models using heterogeneous training backends for student and teacher models.
Survey introducing reinforcement learning methods to economists for solving high-dimensional dynamic programming problems in economic modeling.
Method automating metadata curation for museum audiovisual archives using multimodal grounding in existing collection databases.
Research using foundation model surrogates with active learning for materials discovery, reducing experimental cycles needed for optimal material identification.
NCCL EP presents a unified expert parallel communication API built on NCCL for GPU-initiated RDMA operations in Mixture-of-Experts LLM architectures.
Deep Adaptive Model-Based Design of Experiments combines deep learning with adaptive sequential design optimization for efficient nonlinear dynamical system parameter estimation.
Study on multi-agent routing architectures identifying how failure propagation differs in tree-like versus cyclic execution graphs for AI reasoning systems.
Research exploring agentic frameworks with domain-specific tools for Verilog code generation, comparing impact versus traditional LLM approaches.
Study analyzing chain-of-thought faithfulness evaluation in LLMs across 12 models, showing measurement methodology significantly affects reported faithfulness percentages.
Research demonstrating mathematical isomorphism between ant colony decision-making and random forest ensemble learning under stochastic ensemble intelligence framework.
TimeTox is an LLM-based pipeline using Google's Gemini to automatically extract time toxicity metrics from clinical trial protocol documents.
GitHub Actions tool for AI agents enabling faster CI feedback loops with mocked runners, caching, and agent-driven test fixes without pushing.
Shopify launches Agentic Storefronts allowing merchants to sell through ChatGPT, Copilot, Google Search, and Gemini via centralized management.
GolfStudent v2: 24M parameter LLM compressed to 15MB using GPTQ-lite quantization and Muon optimizer with efficient architecture.
AegisFlow: Open-source AI gateway in Go providing routing, security policies, rate limiting, cost tracking, and observability for LLM providers.
Overview of AI training projects on Alignerr platform including Prism, Code Human, and Rainforest with details on participation.
AI video ad generator transforming product URLs into structured video ads for e-commerce platforms like Shopify and Amazon.
Post-mortem analysis of failed AI podcast app identifying issues with podcast content discovery and user engagement expectations.
Headless virtual terminal tool enabling AI agents to operate interactive TUI applications without GUI, with example of agent playing NetHack.
AI agent consumer application that autonomously dates on user's behalf via chatting with matches. Tool but limited technical depth.
Go sidecar process manager for cleaning up orphaned stateful processes from Puppeteer/LLMs to prevent memory leaks and OOM crashes.
Discussion on Hacker News about using Claude and LLMs for full code generation in production environments, including challenges with debugging.
Claude Code plugin that aggregates and scores content from 7 social media platforms against user interests, generating HTML dashboards.
Case study building a shared sandbox workspace for two OpenAI agents collaborating via Discord and remote VPS with controlled communication.
Technical guide on multi-agent orchestration using ACPX protocol instead of PTY scraping for communication between coding agents.
OpenAI announces public bug bounty program for identifying safety and abuse risks in AI products.
Technical overview of how Cursor trained Composer 2 using pretraining, RL, and realistic coding benchmarks.
Consensus Code project implementing AI agent coordination through libertarian socialist principles for software development.
Tessera is an open-source framework running 32 OWASP AI security tests against GPT-4o, Claude, Gemini, Llama 3, and other models via CLI.
Bleep is an on-premise AI security proxy that scans text and images for secrets before reaching ChatGPT, supporting agents and custom endpoints.
Llamacpp adds unified system RAM offloading support on Linux for efficient on-device AI inference.