GenGEO: A binary trust registry for AI agent transactions
Machine-readable trust registry for AI agents to verify merchant safety before making autonomous transactions.
Machine-readable trust registry for AI agents to verify merchant safety before making autonomous transactions.
Analysis of LLM capabilities distinguishing crystallized intelligence (pattern learning from training data) from fluid intelligence (novel problem-solving), suggesting LLMs excel at trained domains but lack human-like reasoning.
Open registry of W2A-compatible sensors providing event-driven input streams for AI agents, with install commands.
Article on using AI coding assistants for payments integration, noting limitations in product-specific tasks like webhook verification.
Guide to building and scaling reinforcement learning environments in the era of LLM-based agents.
Interactive experiment using generative UI to create full-page images from search queries, navigating by tapping.
Binary token-native transport protocol for LLM APIs to reduce bandwidth by keeping token IDs as integers instead of converting to UTF-8 JSON.
Wrapper tool converting Claude CLI into OpenAI-compatible API with concurrency, retry logic, and prompt caching.
Developer built compiler to manage AI coding agent skills across multiple platforms (Claude, Cursor, Windsurf) with unified source format.
2026 roadmap addressing AI/ML challenges in smart manufacturing including data management, system integration, and industrial deployment.
AI agent system built on n8n platform for automated ESG performance classification and assessment in European SMEs using expert-validated baselines.
Geometric analysis of emergent misalignment in LLMs through feature superposition, explaining how fine-tuning on narrow tasks induces harmful behaviors.
Clinical chatbot using prioritized evidence RAG with guideline-grounding and verifiable citations to reduce hallucination in medical diagnosis.
Knowledge-driven LLM decision-support system with ontology integration for defect diagnosis and mitigation in laser powder bed fusion manufacturing.
LLM-based intelligent agent platform for stuttering assessment and personalized therapy with clinician-in-the-loop workflow.
Multi-agent system prototype for scientific workflows in hydrodynamics, addressing context limitations of single-agent LLM systems.
LLM-guided evolutionary search establishing new exact Zarankiewicz numbers and bounds through reinforced optimization.
RLHF-based approach for adapting LLMs to match instructor style in automated educational feedback while preserving diagnostic accuracy.
Empirical study showing iterative finetuning on model outputs mostly produces idempotent behavior rather than amplifying tendencies.
Defense system detecting adversarial interaction patterns in LLM agents through low-latency anomaly detection on agent behavior.
Position paper arguing multi-agent safety depends on interaction topology rather than individual model alignment or scale.
Mechanistic analysis of how Llama-3.1-8B performs cyclic reasoning through base-10 addition rather than modular arithmetic.
Neuro-symbolic system combining SNOMED CT ontology with machine learning for interpretable clinical AI predictions.
Benchmark for evaluating process reward models across diverse reasoning tasks beyond mathematics, enabling detection of intermediate reasoning errors.
Framework improving faithfulness in vision-language GUI agents by grounding actions in screen evidence and user instructions via guided advantage estimation.
Position paper proposing agentic systems be designed as token allocation economies with specialized layers for routing, planning, and action selection.
Interactive simulation environment for training multimodal agents to perform Earth observation analysis with tool use and uncertainty resolution.
Neuro-symbolic framework for inducing executable skills from agent interactions, combining LLM reasoning with programmatic logic for long-horizon planning in dynamic environments.
Policy optimization method aligning RL credit assignment with natural reasoning steps in multi-modal tasks at segment granularity rather than token or sequence level.
Study of in-group favoritism biases in persona agents facing contradicting information and methods to mitigate adverse effects on factual accuracy.
Multimodal dataset and recognition framework for non-standard system-level chip design diagrams to improve MLLM understanding of architectural specifications.
Evaluation of cognitive plausibility for computational models of analogy and metaphor including SME, CogSketch, and LLMs using the Minimal Cognitive Grid framework.
Systems framework analyzing AI safety through irreversibility control and deployment friction reduction in autonomous decision-making systems.
Hierarchical tokenization framework for time-series generation enabling user control over temporal granularity from sketches or scratch.
Formal theory of Artificial Jagged Intelligence modeling uneven optimization pressure across capability domains during training as finite-budget gradient allocation.
Method for auditing and composing LoRA adapter libraries with residual merging and reliability assessment for task-level reuse and instance-level selection.
Formal framework for explaining entailments in description logic knowledge bases with user-centered contrast-based approach beyond traditional justifications.
Few-step generative model for offline multi-agent reinforcement learning enabling coordinated inference without sacrificing inter-agent coordination.
Framework grounding multi-hop fact verification in structural causal models using group relative policy optimization to improve LLM reasoning and reduce hallucinations.
Multi-turn legal consultation agent using coverage-driven retrieval control to determine sufficient evidence and relevant legal issues.
Deep research agents for automated scientific discovery using post-training on information-seeking tasks and iterative problem-solving capabilities.
Agent system for human-vehicle collaboration using bidirectional perception and alignment to improve driver-automation coordination and situational awareness.
Analysis of inference scaling strategies including self-consistency, self-refinement, and multi-agent debate for compute-efficient LLM performance improvement.
Evaluation framework for production agentic AI systems addressing compounding errors, tool failures, output drift, and long-horizon task evaluation beyond lab-scale benchmarks.
Multi-agent system using LLMs to translate natural language into constraint programming models with synthesized validation checkers to reduce semantic errors.
Framework for designing latent state representations in world models for agents, categorizing methods by functional purpose rather than implementation approach.
Research on routing mechanisms in AI systems and their impact on trust, cost, quality, and accountability of responses across different service tiers and endpoints.
Study of genre bias in LLM credibility assessment, showing models misclassify entertainment news more than hard news.
Defense mechanism against infectious jailbreak attacks in multi-agent systems through foresight-guided strategies.
Momentum: game with runtime procedural content generation evaluated by autonomous agents for balance and playability.