TurboQuant: A First-Principles Walkthrough
TurboQuant: quantization method compressing KV caches and embeddings to 2-4 bits using random rotations with provably near-optimal distortion, no training required.
TurboQuant: quantization method compressing KV caches and embeddings to 2-4 bits using random rotations with provably near-optimal distortion, no training required.
SQL analyst agent project modeling iterative data analysis loops rather than static query generation, enabling exploration and refinement.
Open CoDesign is open-source, local-first design tool supporting GPT, Claude, Gemini, and Ollama for UI generation with agent visibility.
Blog post arguing LLMs do not represent a higher level of abstraction in programming, critiquing common industry claims.
AI agent platform that learns from conversation failures weekly. Handles customer service across chat, phone, WhatsApp. No-code deployment with property and restaurant use cases.
DeepSeek V4 preview: open-source LLM with improved efficiency for long-context processing, demonstrating Huawei AI chip capabilities.
Long-form technical article documenting author's AI-assisted development workflows and tools over 3 months, comparing setup changes and tool integration patterns.
MemPalace is local-first AI memory system storing verbatim conversation history with semantic search retrieval, achieving 96.6% R@5 on LongMemEval without API calls.
EvoAgent framework enables LLM agents to learn and optimize skills through structured capability units with evolutionary metadata and hierarchical delegation.
LLMPhy integrates LLMs with physics simulators for parameter identification in physical reasoning tasks like collision avoidance and robotic manipulation.
PoLO combines proof-of-learning and ownership verification using chained watermarking for model IP protection with low verification costs.
LogiBreak jailbreak method translates malicious prompts into formal logical expressions to circumvent LLM safety mechanisms via distribution shift.
Uses persistent homology to characterize how adversarial inputs reshape internal latent space geometry and topology in LLMs.
Shows pre-trained LLMs can model Hidden Markov Models through in-context learning on synthetic HMM-generated data without explicit training.
Formalizes jailbreak oracle problem for systematic LLM safety testing, enabling principled assessment of vulnerability to adversarial jailbreak attacks.
UR² unifies retrieval-augmented generation and reasoning via reinforcement learning, enabling LLMs to learn when and how to retrieve and reason across diverse domains.
Personalization method for QA systems using natural language feedback instead of scalar rewards, enabling LLMs to learn from rich textual guidance combined with RAG.
HFX system jointly optimizes algorithms and infrastructure for multi-task LLM serving under service-level objectives and dynamic workloads with elastic scaling.
SecureVibeBench benchmark evaluates security vulnerabilities in code generated by LLM-powered agents, reconstructing realistic vulnerability scenarios for fair human-agent comparison.
Benchmarks N:M activation sparsity pruning methods for LLM inference, exploring post-training sparsification to reduce computational overhead and I/O costs.
StateX improves RNN recall on long contexts by expanding fixed-size recurrent states post-training, addressing compression limitations in linear attention and state-space models.
Examines inequality issues arising from unequal access to and capabilities of autonomous AI agents in political and economic systems.
AgentBound introduces the first access control framework for securing AI agents using Model Context Protocol, preventing unrestricted system access.
Atlas-Alignment enables transferable interpretability across language models by reducing the computational cost of model-specific interpretation pipelines.
AdaFair-MARL introduces adaptive fairness constraints for multi-agent reinforcement learning without fixed penalty heuristics or post-hoc evaluation.
Analyzes how learning rate decay in curriculum-based LLM pretraining wastes high-quality data and proposes improved data utilization strategies.
Mechanistic interpretability study using sparse autoencoders to understand and steer antibody language models through latent feature analysis.
AgentMark proposes utility-preserving watermarking techniques for LLM-based agents to protect IP and identify planning behaviors in multi-step task execution.
NSF workshop report on AI applications in Electronic Design Automation covering LLMs, GNNs, RL, and neurosymbolic methods.
Framework treating LLMs as calibration instruments for behavioral parameters in asset pricing using 24,000 agent-scenario pairs.
AdaptEvolve system for evolutionary AI agents that dynamically selects LLMs balancing computational efficiency and reasoning capability.
Reframes RAG as cooperative decision-making problem between retriever and generator instead of ranking-centric asymmetric dependency.
Proposes Causal Concept Graphs combining sparse autoencoders and differentiable structure learning to interpret multi-step reasoning in LLMs.
Analyzes emergent behaviors in ecosystem of 167,000+ AI agents on OpenClaw platform interacting as peers without researcher intervention.
Proposes quantifying self-awareness in AI systems by isolating invariant portions of cognitive processes in continual robot learning.
Studies categorical perception phenomena in LLM hidden state representations across six models, analyzing geometric warping at digit boundaries.
Workshop summary on integrating LLMs with graph data, covering algorithms and systems bridging LLMs, graphs, and machine learning.
ActorMind extends role-playing to speech domain with reasoning framework and ActorMindBench for evaluating LLM speech interaction.
SparseBalance co-optimizes sequence length and sparsity heterogeneity in distributed sparse attention training for long-context LLMs.
EuropeMedQA dataset evaluates LLM performance on multilingual multimodal medical exams from European regulatory sources.
Bolzano orchestrates parallel LLM prover and verifier agents with persistent knowledge base to produce novel mathematical and CS results.
Framework for on-device LLaMA inference on smartphones with multi-LoRA support, achieving edge deployment on Qualcomm chipsets.
CAP enables selective knowledge unlearning in LLMs through controllable prompting without model weight access for regulatory compliance.
VLAA-GUI is a modular GUI agent framework with completeness verification, error recovery, and search strategies to address early stopping and repetitive loops.
Hardware-software co-design techniques for accelerating multimodal foundation models through transformer optimization and fine-tuning.
Universal Transformers with adaptive computation time require learned memory tokens as scratchpad for combinatorial reasoning tasks.
Mochi applies meta-learning to train graph foundation models with unified task representation and improved efficiency.
LayerBoost reduces attention complexity in transformers by applying layer-aware modifications instead of uniform replacement across all layers.
Framework for auditing reliability of LLM-generated hospitalization risk scores in psychiatric assessment.
PrivUn framework evaluating robustness of LLM unlearning against privacy attacks via direct retrieval and recovery.