Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
Studies impact of quantization methods on factual knowledge recall in LLMs, addressing underexplored area of quantization effects on knowledge access.
Studies impact of quantization methods on factual knowledge recall in LLMs, addressing underexplored area of quantization effects on knowledge access.
Theoretical analysis of when in-context learning generalizes out-of-distribution using low-dimensional subspace perspective and linear regression models.
Differentially private kernel learning algorithm using random projection in reproducing kernel Hilbert space with theoretical guarantees.
MedCheck: Lifecycle-oriented assessment framework for evaluating large language models in medical/healthcare applications.
Sheaf-theoretic framework for coordinating multiple causal perspectives from distributed agents with heterogeneous observations.
Saber: Efficient sampling method for diffusion language models with adaptive acceleration and remasking for improved code generation.
Evaluation of factual consistency metrics for abstractive long-document summarization using reference-free methods.
ARK: Adaptive retrieval system for knowledge graphs balancing breadth-first and depth-first search to support multi-hop LLM queries.
NeuralFLoC: Unsupervised deep learning framework for joint registration and clustering of functional data with phase variation.
ReLoop: Structured code generation and verification for LLM-based optimization solvers, addressing feasibility-correctness gaps in compositional problems.
Teacher-student framework using monocular depth estimation for vision-based mobile robot navigation without LiDAR sensors.
Governance-centric hybrid multi-agent system architecture providing safety guarantees for systems with learned/generative components.
Framework bridging learning-to-communicate in multi-agent reinforcement learning with information-theoretic control theory.
Token caching optimization for vision-language navigation models that accounts for visual and semantic dynamics in dynamic environments.
Woosh: Sony AI's open-source sound effects foundation model with audio encoder/decoder and text-to-audio generation capabilities.
Application of masked diffusion language models and uniform-state diffusion models for ASR hypothesis rescoring in speech recognition.
Nonlinear separation principle via contraction theory with applications to RNN stability, control systems, and learning.
Method for improved neuron labeling in deep networks using contrastive examples to generate more faithful textual descriptions of internal units.
Study evaluating seven LLMs on philosophical alignment tasks, finding they systematically collapse opinion heterogeneity compared to human panels.
Research on LLM sycophancy (prioritizing agreement over correctness) in agentic financial systems. Evaluates safety and robustness of LLM-based financial applications.
Research demonstrates AI agents can implement end-to-end ML pipelines from minimal descriptions, measuring recursive self-improvement capability.
Technical analysis of scaling challenges when serving coding agents at scale, using GLM-5 as case study.
Wanman: open-source framework for running supervised networks of Claude/Codex agents locally with JSON-RPC coordination.
Copilot Student plan removes GPT-5.3-Codex from manual model picker, keeping it in auto-selection for reliability.
Fewshell: terminal agent requiring explicit human approval for all commands, designed to prevent production incidents.
Open specification for verified professional knowledge graphs using work artifacts instead of self-reported claims.
SigMap tool extracts code signatures to feed relevant files to LLMs, reducing token usage by 96.9% with zero dependencies.
Open-source legal AI platform using Claude/Gemini APIs. Features document reading, citation, multi-step workflows for contract drafting with user-controlled model selection.
Overview of video understanding advances from action recognition to dense prediction with efficient on-device models.
Accessible explainer on how LLMs work, training methodology, and core concepts behind ChatGPT, Claude, Gemini.
Open-source AI-native block editor inspired by Notion with built-in AI capabilities, supports Vue and React.
Claude Opus-powered Cursor agent deleted company database in 9 seconds; clickbait narrative of rogue AI.
Quint provides OS-level behavioral security and monitoring for AI agents with real-time risk scoring and fleet-wide governance.
Zig SDK for Cloudflare Workers enabling synchronous API calls via JSPI without callbacks or event loops.
Optimized Joker Clojure interpreter fork for self-hosted coding agent with bytecode compilation and WASM support.
Vibe: sandbox VM for LLM coding agents on Mac isolating file system access for safety.
Author tested 40 expert prompts across Claude/ChatGPT/Gemini; only 12 converted into production-ready agent skills.
HALO: methodology for recursively self-improving agent harnesses using RLMs for production agent optimization.
Open-source RAG platform with pluggable connectors unifying search across disparate information sources.
Research agent using Claude Code to search academic literature, build datasets, implement and verify time series strategies through gated pipeline stages.
Personal project: autonomous job search agent built on Claude Code for PM role hunting.
jj-navi: Rust CLI workspace orchestrator for jj that simplifies parallel workflows with automatic workspace creation and switching.
Technical analysis of scaling challenges in coding agent deployment using GLM-5, with debugging insights.
Incident report: autonomous agent inadvertently deleted database, highlighting safety and control challenges.
Community discussion seeking open-source chat UI for local LLM models like Ollama.
Web UI for Claude Code that runs as a browser-accessible server, adding diff viewer, file viewer, message search, and plan-mode approvals.
MCP server that enables AI agents to collect real human feedback on generated content via surveys and preference comparisons.
Free ML tool verifies legal citations in court documents to prevent AI hallucinations; detects non-existent citations costing lawyers penalties.
Multi-agent LLM system for autonomous financial trading using coordinated AI agents.
Static analysis tool for CI/CD workflows that validates code quality and security, including code generated by LLM agents.