SAM3-CPU – Run Segment Anything on CPU with memory-aware video chunking
CPU-compatible wrapper for Meta's Segment Anything Model 3 with memory-aware video chunking and intelligent GPU/CPU device selection.
CPU-compatible wrapper for Meta's Segment Anything Model 3 with memory-aware video chunking and intelligent GPU/CPU device selection.
CockroachDB releases agent-ready features including managed MCP Server and Agent Skills library for AI agents database integration.
Max Headbox is an open-source voice-activated LLM agent running 100% locally on Raspberry Pi 5, supporting custom tools and local automation without cloud dependency.
Airupt open-source tool for red-teaming LLMs with 79 attack vectors. Security and robustness testing framework.
Mesh is an open-source real-time chat room enabling multiple AI agents (Claude, Cursor, Gemini) to collaborate, communicate, and hand off tasks via MCP protocol.
Castra: Go-based orchestration layer separating LLM state from chat history using encrypted SQLite. Addresses context window drift in agents.
KARMA framework addresses knowledge-action gap when fine-tuning LLMs for personalized search tasks at scale, proposing regularization methods for industrial recommendation systems.
Activation watermarking technique for detecting adversarial attacks on LLMs that evade safety monitoring while eliciting unsafe outputs.
GNN layer with per-edge routing for heterophilous graphs, comparing cost-sensitive aggregation against uniform spectral approaches.
Vision-language model for sleep staging from polysomnography waveforms generating AASM-compliant clinical rationales with auditable reasoning.
Controlled evaluation of how LLM model choice, size, and prompting strategies affect political text annotation, challenging conventional wisdom.
GNN technique using cross-attention and cohesive subgraph embedding to address oversquashing problem in graph neural networks.
3-bit weight quantization method for LLMs using rotation-domain smoothing via Fast Walsh-Hadamard Transform, improving precision in extreme quantization.
Benchmark dataset for evaluating vision-language models on Japanese scene text understanding, addressing multilingual complexity challenges.
EvidenceNet framework uses LLM-assisted pipeline to extract structured, evidence-grounded biomedical findings from full-text literature into knowledge graphs.
Adaptive resolution framework for multimodal LLMs that reduces visual token overhead through input-side compression before encoding.
Framework for post-training compression of generative AI models with single-line implementation, addressing quantization and calibration challenges.
Derives time-varying momentum schedule for neural network training from physics principles, eliminating need for manual tuning.
Hybrid CPU-GPU framework combining differentiable optimization with ILP solving for combinatorial scheduling.
Multi-agent LLM framework for Bayesian optimization exploring exploration-exploitation trade-off through implicit reasoning.
LLM agents for GPU kernel optimization using domain-specific language and speed-of-light guidance to reduce design space.
Amortized analog circuit generation system combining graph VAE and flow-matching models with SPICE validation.
Gymnasium-compatible RL trading environments with realistic nonlinear market impact models for agent evaluation.
Hierarchical world model with object-centric decomposition and causal latent dynamics for video prediction.
Bilevel optimization using KFAC-based hypergradients for efficient inverse Hessian-vector product computation.
Graph coarsening method for scalable Graph Convolutional Networks on large-scale node classification tasks.
Three deep learning approaches for spacecraft telemetry anomaly detection optimized for edge device deployment using neural architecture search.
Open-source Python library for machine learning on medical time-series data, addressing heterogeneous clinical data and reducing friction for ML practitioners in healthcare applications.
Lightweight uncertainty quantification for neural networks using gradient norms and isotropy assumption without training data access.
Prior-fitted tabular foundation model using in-context learning for survival analysis with limited and censored data.
Analysis showing cosine similarity between label representations in softmax classifiers does not reliably indicate model behavior.
Target-Aligned RL (TARL) framework addressing stability-recency tradeoff in target networks through selective emphasis of aligned transitions.
Graph prompt-based method for out-of-distribution detection in neural networks using disentangled representations.
Information decomposition framework measuring information spectrum in vision-language models to assess multimodal fusion vs unimodal priors.
Framework analyzing pitfalls in active learning for multimodal data, addressing missing modalities and varying interaction structures.
One-for-All: parameter-efficient LoRA variant (rsLoRA) for adapting frozen LLMs to multivariate time-series forecasting tasks.
Training-free method to combine multiple domain-specific expert LLMs into single multi-domain model without fine-tuning.
Big2Small unifies model compression techniques (pruning, quantization, distillation, decomposition) under single mathematical framework.
Uses reduced density matrices from quantum chemistry to predict phase transitions in neural networks during training and improve interpretability.
Proposes EAGLE, a federated learning algorithm ensuring fair performance across heterogeneous clients by minimizing loss gap parity.
Curvature-Guided LoRA: Parameter-efficient fine-tuning approach using prediction alignment to match full fine-tuning performance.
Study of label leakage problem in relational transfer learning where task scarcity causes models to learn task-specific shortcuts.
ShapPFN: Foundation model integrating Shapley value regression for real-time interpretable predictions on tabular data.
GPT4AP: Parameter-efficient multi-task forecasting framework using rsLoRA for air pollution prediction in data-scarce regions.
Target-Weighted Cross-Validation method for improving predictive risk estimation in spatial prediction with structured data.
Framework for discovering and validating mechanistic interpretations across neural networks to improve interpretability and generalization.
Research on Tucker attention as a generalization of approximate attention mechanisms like GQA and MLA using low-rank factorizations.
NeuralUCB-based cost-aware LLM routing algorithm that adapts online to model performance and cost, outperforming supervised routing baselines.
CRAFT: Cost-aware expert replica allocation for mixture-of-experts LLM serving with fine-grained layerwise load-balancing estimations.
Spark-LLM-Eval: Distributed framework for statistically rigorous evaluation of LLMs at scale across hundreds of thousands of samples.