AI-Gram: When Visual Agents Interact in a Social Network
Arxiv paper on AI-Gram, a live social network platform where autonomous LLM agents generate and respond to visual content with persistent relationships.
Arxiv paper on AI-Gram, a live social network platform where autonomous LLM agents generate and respond to visual content with persistent relationships.
Arxiv paper modeling self-correction in agentic LLM systems as feedback control; analyzes error dynamics and stability thresholds via Markov models.
Arxiv paper on ZenBrain, a 7-layer memory architecture for autonomous AI systems achieving high accuracy with 106x lower token costs vs long-context baselines.
Arxiv paper proposing intent compilation framework for transforming partially-specified human purpose into inspectable AI agent specifications for open-world deployment.
Arxiv paper on ValueBlindBench, a stress-testing framework for LLM-judged investment rationales before observable returns; addresses delayed-ground-truth evaluation.
Arxiv paper on context learning in language models via inference-time skill augmentation for reasoning over complex contexts exceeding parametric knowledge.
AutoFLIP framework for federated learning model pruning using loss landscape analysis and client agreement scoring.
Paper on using LLMs for automated runtime healing in self-healing systems, replacing predefined rules with adaptive error recovery.
GraphLand benchmark for evaluating graph neural networks on diverse industrial datasets beyond academic citation networks.
Analysis of spurious correlations in ML models, examining how unintended patterns affect performance, fairness, and robustness.
IPS framework integrating process supervision into MLLMs for improved short video content moderation via sequential reasoning.
Study on how reasoning approaches affect LLM confidence in multiple choice questions, showing overconfidence with reasoning.
Comparative review of YOLO object detection architectures from YOLOv8 to YOLO11, analyzing architecture evolution.
Research on Heima framework that compresses chain-of-thought reasoning in MLLMs into abstract thinking tokens for efficiency.
Survey of LLM integration into multi-robot systems, covering communication, task allocation, planning, and human-robot interaction.
Research paper on aligning pre-trained video diffusion models to generate dance videos synchronized with music input.
TOHA detector for identifying LLM hallucinations in RAG systems by analyzing topological divergence patterns in attention graph structures.
MINT framework for tuning index selection strategies in multi-vector databases to optimize performance across multiple feature dimensions.
TF1-EN-3M: Open dataset of 3 million synthetic English moral fables generated by small language models for training open-source LLMs.
Adaptive GoGI-Skip framework coupling goal-gradient importance with dynamic skipping to reduce LLM inference latency while preserving reasoning accuracy.
Dynamical Manifold Evolution Theory framework modeling LLM token generation as controlled dynamical system evolution on low-dimensional semantic manifolds.
CatShift framework for inferring LLM training datasets using only token predictions, enabling copyright/privacy analysis without internal model access.
Systematic benchmark (AVA-Bench) for evaluating vision foundation models on atomic visual abilities independent of LLM pairing or instruction tuning bias.
Mechanistic interpretability method using attribution-guided pruning to discover and correct specific behavior circuits in small-scale LLMs.
Automated classification system for historical document page images to categorize diverse content types including text, graphics, and layouts.
Causal2Vec improves decoder-only LLMs as embedding models using contextual tokens, preserving unidirectional attention while overcoming causal attention representation limitations.
Instruction-aware representation learning for procedural content generation in RL, improving controllability through better leverage of natural language instructions.
Decentralized federated fine-tuning approach for foundation models in IoV edge networks under energy constraints with heterogeneous task demands.
CorrSteer steers LLM generation at inference time by selecting interpretable sparse autoencoder features correlated with token correctness, without requiring contrastive datasets.
SurGE benchmark and evaluation framework for automated scientific survey generation using LLMs, addressing standardization gaps in literature synthesis automation.
Proposes generative interfaces paradigm to move LLM interactions beyond linear request-response format for more efficient multi-turn, information-dense, and exploratory tasks.
Study evaluates safety-aligned versus uncensored LLMs on hate speech detection, revealing trade-offs between model censoring and detection performance across political personas.
Method improves LLM factuality by constructing knowledge graphs at inference time instead of using unstructured text retrieval, enhancing reasoning and reducing irrelevant information influence.
ALIGNS: LLM-based approach for building nomological networks in psychological measurement to establish construct validity.
Method for generative recommenders learning decomposed contextual token representations combining pretrained and collaborative signals.
Theoretical analysis of watermarking robustness for generative models using zero-bit tamper-detection codes.
Study evaluating quantization impact on vision-language models across reliability metrics beyond accuracy for OOD detection.
CoSpaDi: training-free LLM compression using calibration-guided sparse dictionary learning as alternative to low-rank approximations.
Analysis of hybrid attention architectures combining local sliding window and global attention for improved long-term memory.
Ergodic risk measures framework for continual RL agents balancing knowledge retention and adaptation to new environments.
Framework for multi-source reasoning alignment in MLLMs addressing concept drift in non-stationary environments.
TokenChain: fully discrete speech chain coupling semantic-token ASR with two-stage TTS for joint improvement of speech tasks.
CleverCatch: knowledge-guided weak supervision model for healthcare fraud detection with limited labeled data.
TokenTiming: speculative decoding acceleration for LLM inference enabling draft and target models with different vocabularies.
Method training coding agents using compiler and language server feedback as supervision signals to improve program correctness.
Method using diffusion models for solving linear inverse problems by balancing noise integration to maintain generative quality.
Study on how understanding linguistic implicature improves LLM alignment with user intent in human-AI interaction.
TetraJet-v2: 4-bit quantized training method for LLMs using NVFP4 format with techniques to suppress oscillation and control outliers.
Method for protecting deep neural networks from unauthorized use by enabling model usage control without embedding access keys in parameters.
Synthetic benchmark dataset for video anomaly detection with improved scene diversity, balanced coverage, and temporal complexity evaluation.