Q: A Slim LLM CLI; For asking and debugging from cli directly
Q: slim shell-based LLM CLI for terminal queries with session logging and secret redaction.
Q: slim shell-based LLM CLI for terminal queries with session logging and secret redaction.
AgentCheck: pytest-style testing framework for AI agent behavior validation and regression testing.
NPM package providing 18 security skill packs for reviewing and hardening LLM applications.
IBM Bob: AI development partner covering full software development lifecycle with enterprise governance and security controls.
Rust Bucket: Agent-first Rust project bootstrapper for rapid project initialization.
Bug bounty program offering up to $2,500 for documenting real-world AI misalignment cases and safety failures.
Analysis of AI-driven tech layoffs and workforce displacement, contrasting Silicon Valley narrative with broader corporate adoption.
Mistral Workflows: enterprise orchestration layer for AI agents built on Temporal, enabling production-ready durable AI processes.
Perplexity research on search-augmented LLMs for improving frontier AI model accuracy in knowledge-intensive tasks.
Informational RFC proposing Prompt Transport Protocol for transmitting intents and hints for local generative model reconstruction.
Browser-based operating system with integrated AI agent (DAEMON) capable of building apps and managing files autonomously.
Conceptual architecture for AI systems based on version-controlled knowledge base with provenance and rights management.
SketchVLM: framework enabling vision-language models to produce editable SVG overlays explaining visual reasoning on input images.
Coregit provides versioned filesystem for AI agents with git-based commits, forks, and time-travel snapshots via REST API.
Analysis measuring how much agent harnesses versus LLMs contribute to planning agent performance on fixed models.
SIV-Bench video benchmark for evaluating multimodal large language models on social interaction understanding and reasoning tasks.
Research on internal mechanisms of vision-language models for anomaly detection, identifying sparse neurons without external adapters.
Statistical physics analysis of random feature models investigating training and test error beyond mean kernel approximation.
SQLyzr platform for fine-grained evaluation and analysis of text-to-SQL models built on LLMs, providing detailed insights beyond aggregate benchmark scores.
Deep learning method for automated congenital heart disease detection from phonocardiograms using feature fusion.
Physics-Informed Neural Networks compared with numerical methods for analyzing bending behavior in perforated nanobeams.
Liquid Neural Networks applied to natural gas spot price forecasting, addressing nonlinear dynamics and regime changes in time-series prediction.
Graph-conditioned trust-region method reduces query costs for low-depth Quantum Approximate Optimization Algorithm implementations.
Intrinsic Mutual Information used as modulator for preference optimization in Large Language Models with reduced hyperparameter tuning.
Energy-aware neural architecture design evaluated across 2,203 experiments, optimizing for computational cost alongside accuracy in ML models.
Nautile-370M is a 371M-parameter language model combining spectral memory and attention for efficient reasoning with strict parameter budgets.
Comparative analysis of Upper Confidence Bound algorithms in Adaptive Deep Neural Networks for energy-efficient edge computing inference.
Time-varying graph neural ODE framework for dynamic graph representation learning with adaptive message passing mechanisms.
Method to estimate closed-source LLM parameter counts by measuring factual capacity, providing tighter bounds than inference economics.
Study of masked diffusion language models with blockwise locality, comparing training stability and structured generation against autoregressive LLMs.
Depth pruning for LLMs that treats layer redundancy functionally relative to evaluation objectives rather than as inherent structural property.
Nemotron 3 Nano Omni: open multimodal model supporting audio, text, images, and video with improvements in document understanding and localization.
Compute Aligned Training method that aligns LLM post-training objectives with test-time inference procedures that aggregate multiple outputs.
CoreFlow: geometry-preserving low-rank flow model for learning matrix-valued distributions from high-dimensional incomplete data.
Odysseys benchmark for evaluating web agents on realistic long-horizon multi-site tasks requiring cross-domain reasoning and sustained context.
PolyKV system enabling multiple concurrent LLM agents to share compressed KV cache pool, reducing memory overhead for multi-agent inference.
SWIFT framework that transfers learned workflow topologies across tasks instead of searching per-task, reducing computational cost for agentic workflow design.
Feasible-first exploration for ML model deployment optimization handling hierarchical search spaces with invalid configurations.
Zero-shot coordination training for multi-agent RL to cooperate with unknown agents in sparse reward tasks.
Position paper arguing knowledge distillation must evaluate preserved teacher capabilities beyond task metrics for reliable deployment.
Low-rank adaptation approach for unified multi-task EEG analysis with self-supervised pre-trained models.
Data cleaning method for tabular foundation models addressing missing values, outliers, and duplicates to improve zero-shot accuracy.
Analyzes reliability of Vision-Language Models as evaluation judges using conformal prediction to provide calibrated uncertainty intervals.
DGLight framework combines DQN critic with reinforcement learning to fine-tune LLMs for traffic signal control optimization.
Integer-only quantization approach for FlashAttention in vision transformers addressing scale explosion and numerical stability challenges.
VAE-based two-stage framework for imbalanced classification combining generative and discriminative modeling.
GNN-based approach for imputing missing modalities in patchwork learning where different clients have varying modality availability.
Safe RL method preventing unsafe state visitation during training via Q-learning with strict safety constraints.
Analyzes limitations of epistemic uncertainty quantification in latent space models for model-based RL with image observations.
Fisher-guided token quantization for efficient federated fine-tuning of LLMs on edge devices with heterogeneous bandwidth constraints.