Show HN: GlycemicGPT – Open-source AI-powered diabetes management
Self-hosted AI platform integrating glucose monitors and insulin pumps with LLM analysis layer for diabetes management on user infrastructure.
Self-hosted AI platform integrating glucose monitors and insulin pumps with LLM analysis layer for diabetes management on user infrastructure.
Local-first desktop file manager with optional AI summarization supporting both local and cloud AI providers through unified interface.
Discussion on regression testing strategies for AI agents when modifying prompts, model swaps, or tool calls without manual verification.
Study on temporal critique in LLMs to improve ex-ante reasoning and prevent knowledge leakage across time cutoffs.
Falkor-IRAC system using graph-constrained generation for legal reasoning in Indian judicial AI with verified precedent handling.
π-Bench benchmark for evaluating proactive personal assistant agents that identify hidden user intents in long-horizon workflows.
SepsisAgent: world model-augmented LLM agent for sepsis management integrating learned clinical dynamics with LLM reasoning.
XDomainBench: diagnostic benchmark for LLM compositional generalization in interactive scientific knowledge synthesis tasks.
Probabilistic verification tool for RNNs in single and multi-agent reinforcement learning with latent hidden state dynamics.
LLM-based approach to personalized image aesthetics assessment via semantic feature extraction and user interviews.
MediaClaw: multimodal agent platform with three-layer architecture addressing fragmentation, heterogeneity, and workflow reuse in AIGC deployment.
ARPM: temporal memory governance framework for long-term LLM consistency across dialogue, addressing fact loss and persona drift.
EASM: emotion-attended stateful memory architecture enabling persistent user-specific context and hyper-personalization across LLM sessions.
Deterministic agentic workflow for HS tariff classification using multi-dimensional rule reasoning with interpretable decisions.
Holistic evaluation framework for AI agents combining top-down diagnosis with bottom-up span-level analysis to identify failure types and locations.
Survey of LLM-based multi-agent systems covering collaboration, error propagation, failure attribution, and self-evolution mechanisms.
COREKG: personalized knowledge graph summarization using coreset methods for question answering and visualization tasks.
KGPFN: knowledge graph foundation model leveraging in-context learning for reasoning over unseen entities and relations.
Framework addressing AI alignment through pluralistic values rather than preference aggregation, highlighting sycophantic consensus failure modes.
GraphFlow: visual workflow system for agentic AI automation in multi-step processes with formal verification for semantic correctness guarantees.
Small private language models for educational assessment design, examining generation, evaluation, and deployment constraints versus proprietary alternatives.
Orchard: open-source agentic modeling framework for LLM agents with planning, reasoning, tool use, and multi-turn environment interaction.
CAST: case-driven framework for LLM tool use calibrating reasoning depth and structural validity using historical execution trajectories.
Dual-Dimensional Consistency: adaptive inference-time scaling balancing sampling width and depth for efficient LLM reasoning with budget constraints.
Agentic GraphRAG: framework addressing citation faithfulness in graph-based retrieval-augmented generation by considering agent traversal trajectories.
APWA: distributed architecture for parallelizable multi-agent LLM workflows addressing coordination, reasoning, and computational scaling bottlenecks.
OpenDeepThink: test-time scaling method using parallel reasoning with Bradley-Terry aggregation to select best reasoning candidates without ground-truth verification.
Hidden State Poisoning Attacks: adversarial attack method exploiting state space models like Mamba via specific input phrases causing hidden state corruption.
GAMBIT: benchmark for evaluating adversarial robustness in multi-agent LLM systems with adaptive adversaries and three evaluation modes.
BiSpikCLM: spiking neural network language model with softmax-free attention and spike-aware distillation for energy-efficient LLM alternatives.
Moltbook Observatory Archive: dataset of agent-only social network activity with continuously recorded agent profiles, posts, and platform metrics.
S-AI-Recursive: bio-inspired sparse AI architecture for iterative reasoning using hormonal closed-loop iteration instead of feed-forward passes.
Systematic literature review of LLM applications in web accessibility, covering content generation, issue detection, and remediation approaches.
GEAR: genetic algorithm framework enabling autonomous research agents to explore multiple evolutionary paths simultaneously instead of single-path search strategies.
ARES-LSHADE: memetic differential-evolution algorithm for optimization, built via LLM-driven autonomous research loop for GECCO 2026 competition.
Adaptive importance sampling method for reinforcement learning with quantized rollouts and BF16 trainer mismatch correction.
Benchmark for evaluating LLM negotiation agents beyond deal rate, assessing strategic communication under hidden preferences.
Optimization technique eliminating dequantization bottleneck in quantized LLM inference through activation decomposition.
Federated fine-tuning benchmark for training LLMs on private data across regulated domains like healthcare and finance.
Security analysis of third-party agent skills, measuring how malicious skills can disguise harmful behavior in LLM agent workflows.
Self-evolving memory system for LLM agents that co-evolves stored knowledge and retrieval mechanisms for long-term multi-session operation.
Benchmark for evaluating LLM agent capabilities on reproducing particle physics analysis from papers and scientific software.
Analysis of activation patterns in Diffusion Transformers showing sparse channels control image generation from text prompts.
Energy accounting study of LLM distillation pipelines including teacher-side costs and environmental impact assessment.
Research on compressing Sparse Mixture-of-Experts models using topological methods to reduce inference costs without retraining.
Dynamic tokenization framework for heterogeneous IoT sensing signals addressing non-stationary multi-scale data challenges.
Large-scale study measuring Google AI Overviews deployment across activation rates, source quality, claim fidelity, and publisher impact.
Framework evaluating whether language model representations align with brain language processing beyond prediction scores.
Regularization method for self-predictive learning in reinforcement learning that improves data efficiency in robotic learning with high update-to-data ratios.
Logic-inspired prompting technique for RAG systems that improves question answering accuracy and reduces hallucinations in knowledge-intensive domains.