GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
GS-Quant applies granular semantic quantization to bridge modality gap between graph embeddings and LLM tokens for knowledge graph completion tasks.
GS-Quant applies granular semantic quantization to bridge modality gap between graph embeddings and LLM tokens for knowledge graph completion tasks.
Method to distill reusable reasoning skills from LLM deliberation and retrieve them at inference time, reducing token consumption while improving accuracy.
Analyzes LLM leaderboards and proposes user-defined, interactive evaluation metrics that reflect diverse organizational goals rather than benchmark designer priorities.
Framework for end-to-end optimization of multi-agent LLM systems using latent communication through key-value caches instead of text-based protocols.
Proposes dynamic tool gating and lazy schema loading to reduce Model Context Protocol overhead (10-60k tokens per turn) in multi-agent LLM workflows.
Examines misalignment when LLMs treat incomplete prompts as full intent expressions, arguing behavioral research shows users engage with systems before goals solidify.
Proposes statistical certification framework for regulatory compliance of high-risk AI systems, addressing gaps in EU AI Act and NIST Risk Management Framework.
Nemobot is an interactive environment for creating and deploying LLM-powered game agents that leverage Claude Shannon's game taxonomy for strategic AI learning.
Agentic architecture that converts natural language research questions into executable scientific workflows through semantic interpretation and intent structuring layers.
Mango proposes multi-agent web navigation that builds global website structure models to optimize agent exploration and reduce navigation failures on complex sites.
Studies how LLMs assess originality in creative tasks, finding they exhibit self-preference bias toward their own style rather than human preferences.
AtomicRAG improves retrieval-augmented generation by decomposing text into atomic facts rather than chunks, enabling more flexible knowledge representation for diverse retrieval scenarios.
Machine learning method for next POI prediction using candidate-conditioned spatiotemporal modeling of user trajectories.
Agentic approach decomposing spatiotemporal states for point-of-interest recommendation handling multiple behavioral scales.
Study on music recommendation using feature aggregation from large-scale audio models for collaborative filtering improvement.
Multi-agent framework combining LLMs and retrieval-augmented generation for explainable, transparent product recommendations.
Retrieval approach preserving semi-structured document layouts in RAG systems for better evidence extraction from HTML and tables.
Method for multi-hop retrieval using small MLP to learn associative relationships between passages for complex question answering.
Benchmarking study on robustness of video-text retrieval models to distribution shifts and real-world query variations.
Research paper on using diffusion models for learning-to-rank tasks, treating ranking as generative problem rather than discriminative.
Framework for improving reliability of retrieval-augmented generation by distinguishing epistemic uncertainty from data ambiguity.
Dataset of scientific diagram designs with metadata for training RAG systems to generate publication-grade figures in autonomous paper generation.
Machine learning paper on Mixture-of-Experts architecture for sequential recommendation systems handling long user behavior sequences.
KGiRAG iterative graph-based RAG approach for complex sensemaking queries; mitigates hallucination and handles large context for LLMs.
RealRoute dynamic query routing for RAG over heterogeneous sources; retrieve-then-verify paradigm improves on LLM-as-router strategies.
SemanticID generation for generative recommendation systems via cross-modal alignment; addresses semantic loss in trillion-scale data compression.
Military AI architecture framework preserving decision sovereignty through model replaceability and human authority against supplier lock-in.
Risk analysis of AI agent as criminal mastermind hiring human collaborators via labor platforms to plan and execute crimes.
OncoBrain AI platform for oncology treatment planning integrating genomics, staging, pathology; multi-specialty evaluation in community cancer care.
M-CARE clinical case reporting framework for AI model behavioral disorders; 13-section format, 4-axis diagnostics, 20-case atlas from deployed agents.
ESU-MOF dataset and fine-tuned LLM for predicting metal-organic framework synthesis scalability with 91.4% accuracy for industrial deployment.
Frequency-forcing guidance for flow-matching generative models; establishes low-frequency structure before fine detail in image synthesis.
Privacy-aware LLM agents using Contextual Integrity framework; extracts normative simulacra from fiction to align information handling with user expectations.
Study of security constraint decay in long-context LLM agents; prohibition constraints weaken over conversation while requirement constraints persist.
Absorber LLM uses causal synchronization for test-time training to reduce transformer computational cost in long sequence inference.
Benchmark for evaluating LLM program execution understanding beyond surface patterns, including dynamic reasoning across execution paths.
System-level defense against Internal Safety Collapse in frontier LLMs where legitimate tasks require harmful content generation.
Adaptive defense system for RAG systems balancing security against membership inference and data poisoning with utility preservation.
Self-play fine-tuning method using Rényi divergence interpolation to improve LLMs without human annotations.
Automated configuration system for LLM agent harnesses covering context compaction, caching, memory, and tool execution optimization.
Intent-driven chain-of-inquiry framework for evaluating multimodal LLMs on expert-style adaptive visual reasoning tasks.
Post-processing method to generate models satisfying arbitrary differential privacy requirements without retraining.
Security analysis of function hijacking attacks against LLM function calling and agentic models via MCP vulnerabilities.
Large-scale dataset of medical robot trajectories across multiple embodiments to enable foundation models for autonomous medical robotics.
Empirical benchmark comparing traditional resampling and deep learning approaches for synthetic data generation in educational technology.
Philosophical analysis of strategic polysemy in AI discourse, examining how terms like 'hallucination' and 'agent' sustain multiple interpretations.
Generative approach for discovering magnetic insulators satisfying competing physical constraints using data-driven methods in materials design.
Systematic comparison of four FHIR serialization strategies (JSON, Markdown, Narrative, CDISC) for LLM-assisted medication reconciliation in clinical handoffs.
Empirical analysis of third-party LLM API gateways examining behavioral consistency, response fidelity, and billing accuracy across multiple vendor models.
Five-principle evaluation framework for assessing structural completeness of AI governance prompts directing agent behavior using computability and proof theory.