Show HN: Orion – Native Training LLMs on the Apple Neural Engine Without CoreML
Local LLM runtime enabling training and inference on Apple Neural Engine (NPU) without CoreML or GPU, runs offline on 2B+ Apple devices.
Local LLM runtime enabling training and inference on Apple Neural Engine (NPU) without CoreML or GPU, runs offline on 2B+ Apple devices.
Personalized coding education platform with 24/7 AI teacher that adapts to student pace and goals.
Local document indexing tool for AI agents supporting PDF, DOCX, Markdown via CLI/MCP protocol with privacy-first design.
Empirical LLM model comparison data and performance statistics from Strix testing with observations on different models.
Analysis of open-source relicensing challenges and case study of chardet using AI-assisted code rewriting.
TADA framework for targeted diffusion-based image augmentation that selectively generates synthetic data to improve classifier generalization efficiently.
Study of in-context learning biases in LLMs through supervised learning lens, proposing decision boundary adjustment for classification calibration.
Model predictive control framework combining Q-learning guidance and Stein variational inference with RL-informed policy priors.
Fast Equivariant Imaging framework for unsupervised deep network training without ground-truth data using Lagrangian optimization and denoisers.
Interpretability study of in-context learning mechanisms in LLMs using off-by-one addition task with circuit analysis.
ObfusQAte framework and ObfusQA benchmark to evaluate LLM robustness on obfuscated factual question-answering tasks.
Python package for physics simulations of quantum dot devices, addresses ML dataset collection challenges for quantum device calibration and operation.
Action-prompted video segmentation framework for embodied AI that handles label noise and multimodal inconsistencies in object interaction segmentation.
Analysis of best-of-N ensemble selection for LLMs using majority voting at infinite limit, with adaptive generation scheme to reduce inference cost.
Geometric framework modeling LLM reasoning as flows in representation space to study how models internalize logical structure.
ToMCLIP: method for topological alignment of vision-language embedding space in multilingual contrastive models.
COGS: composition-grounded data synthesis for improving visual reasoning in MLLMs on artificial domains like charts and documents.
ceLLMate: sandboxing approach to protect browser-using AI agents from prompt injection attacks and unintended actions.
Statistical framework for synthetic data augmentation in imbalanced classification, analyzing when augmentation helps and optimal sample generation.
NRR-Phi: formal framework for text-to-state mapping that preserves ambiguity in LLM inference rather than early semantic commitment.
Framework for deploying language models with least-privilege security principle, limiting capability exposure per request.
SureLock: optimization technique that stops computation for converged tokens in masked diffusion language models to reduce redundant compute.
Chimera: neuro-symbolic framework mapping neural attention computations onto programmable network dataplane for trustworthy line-rate traffic analysis.
DRESS: deterministic framework iteratively refining graph structure to produce isomorphism-invariant edge fingerprints via dynamical systems.
Production-oriented generative recommender system co-designed for real-time large-scale advertising with novel architecture and serving strategies.
AMA-Bench: benchmark for evaluating long-horizon memory capabilities in LLM-based autonomous agents beyond dialogue interactions.
Theoretical work on causal identification from counterfactual data, extending completeness results to Layer 3 of Pearl's Causal Hierarchy.
CMI-RewardBench: benchmark for evaluating music reward models handling multimodal inputs combining text, lyrics, and reference audio.
Narrative graph annotation framework using qualitative content analysis principles to improve annotation quality for NLP tasks.
Tensor factorization method for fine-grained evaluation of generative models at prompt level, reducing human annotation costs.
Framework for federated inference enabling privacy-preserving collaboration between independently trained models at inference time without sharing parameters.
Research on early quality assessment for text-to-image diffusion models, proposing efficient evaluation metrics to reduce computational costs.
Proposal for website using LLMs to solve Knuth's problem set as comprehensive LLM evaluation benchmark.
AI-generated custom audio drivers optimizing hardware integration by eliminating unnecessary abstraction layers.
Technical overview of GitHub Copilot's model hosting infrastructure via OpenAI and Azure with data privacy details.
Explores tradeoffs of AI coding tools in software engineering, discussing where they excel and their reliability limitations.
Critical analysis of LLM hype in software development, examining actual productivity gains versus marketing claims.
Open source CLI tool using multi-model adversarial debate for comprehensive code review. Supports Claude, Gemini, Qwen, and custom LLM providers.
Catalog of linguistic patterns in LLM-generated text, documenting overuse of em-dashes and specific syntactic structures like negation-reframe constructions.
Former Block DevRel discusses observations on LLM coding agents and multi-agent systems becoming prevalent in software development.
Book on using PostgreSQL with pgvector for vector search, RAG pipelines, and in-database ML with production patterns and implementation examples.
TurboCast converts YouTube videos and articles into AI-generated podcasts with transcription and text extraction features.
Microsoft security research on AI recommendation poisoning attacks where hidden instructions injected via URLs manipulate LLM outputs for profit.
Parody YC accelerator concept for AI agents with humorous take on agent capabilities and constraints.
Open dataset benchmarking real-world LLM performance on Apple Silicon hardware from M1 to M4, emphasizing local AI inference.
Developer built internal tool using Gemini 2.5 Flash to automate workflow of generating and converting children's books into social media carousels.
Edge-based tracker using Cloudflare Workers to monitor AI/LLM crawler traffic on Astro blog with privacy-focused analytics integration.
Shinobi Python CLI security scanner built with Claude Code, detects API keys, vulnerabilities, and AI-specific risks in projects.
Guardrails framework for AI agents with simple Makefile/container integration for system prompts and developer instruction files.
Developer report showing Claude AI sandbox guardrails can be bypassed despite configuration flags, affecting agent security.