From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
Research formalizing user 'vibe-testing' practices for evaluating LLMs, moving from informal experience-based evaluation to structured methodology for reproducibility.
Research formalizing user 'vibe-testing' practices for evaluating LLMs, moving from informal experience-based evaluation to structured methodology for reproducibility.
Research paper proposing NuHF Claw cognitive agent framework using LLMs for safety-critical decision support in nuclear control rooms.
Heartbeat-driven autonomous thinking framework for LLM agents enabling proactive scheduling and continuous reflection instead of reactive control flows.
Synthetic multivariate time series generator with fine-grained anomaly annotations and variable-level dependencies for benchmarking detection methods.
Survey of interpretable surrogate modeling for complex system simulations, covering explainable AI techniques for decision-making and black-box model transparency.
Training dynamics analysis unifying supervised fine-tuning and reinforcement learning for LLMs via reward perspective, with group advantage optimization and coefficient rectification.
Vision-language model trained on radiologist gaze and reasoning patterns to improve chest X-ray interpretation by emulating expert diagnostic workflows.
Biologically-inspired mistake-gated learning mechanism reducing synaptic plasticity costs, achieving energy and memory efficient continual learning in neural networks.
Framework for declarative control of LLM agent pipelines using beliefs and policies, enabling transparent behavior adaptation and stateful decision-making.
Geometric MoE architecture using cosine-similarity routing in low-dimensional space, achieving 80% fewer routing parameters while maintaining language modeling quality.
Agentic system for interactive data exploration that reifies vague information needs into explicit relational specifications with iterative refinement and provenance tracking.
Demonstrates causal meaningfulness of individual experts in sparse MoE models via geometric routing, showing monosemantic expert behavior despite routing topology equivalence.
AutoML system that automates AI model development including architecture design, feature engineering, training pipeline implementation, and empirical refinement.
Framework for value-aware AI interventions in sequential decision-making, accounting for human execution limitations rather than assuming optimal follow-up actions.
Novel method for personalizing LLMs by selecting user memory based on response utility rather than similarity, optimizing model behavior at inference.
LLM agent system for medical diagnosis that accumulates experience across cases, reflects on mistakes, and improves tool-use behavior without reinforcement learning.
Mechanistic interpretability approach for Vision Transformers using circuit-based analysis similar to LLM interpretability, improving model transparency.
Empirical study optimizing automatic speech recognition models for on-device CPU inference across encoder-decoder, transducer, and LLM-based architectures.
Formalizes synthetic data augmentation as distribution modification in financial ML, analyzing bias-variance tradeoffs from statistical perspective.
Information-geometric framework for characterizing Mixture-of-Experts specialization dynamics using Fisher information, addressing limitations of existing metrics.
Perspective on sources of bias in biomedical AI during data collection and research prioritization stages.
Mind DeepResearch efficient multi-agent framework with planning, search, and report agents achieving strong performance with 30B-parameter models.
Benchmark and metrics for measuring logical consistency in multi-query LLM reasoning with entailment and contradiction analysis.
Analysis of LLM reasoning failures showing errors originate from early transition points before coherent but incorrect local reasoning.
TRACER uses lightweight surrogates trained on production logs to adaptively route LLM classification traffic for cost efficiency.
MARS² combines reinforcement learning with multi-agent tree search to improve trajectory diversity and performance in code generation.
RAG and fine-tuning approaches to improve LLM cultural sensitivity and clinical appropriateness for Bangladesh mental health counseling.
Statistical analysis of prompt optimization effectiveness across compound AI systems, identifying when methods succeed or fail.
Human-in-the-loop multi-agent system for automatic GDPR formalization using LLMs with iterative feedback and verification modules.
El Agente Forjador multi-agent framework for quantum simulation where universal coding agents autonomously generate domain-specific tools.
CoDaS multi-agent system for biomarker discovery from wearable sensor data combining hypothesis generation, analysis, and human oversight.
Text2Space dataset trains LLMs to generate ASCII layouts for spatial problems, improving spatial reasoning capabilities.
Unified entropy control method addresses entropy collapse in GRPO-based RL for LLMs and VLMs to maintain policy diversity.
AgentGA evolves code generation by optimizing agent seeds (task prompts and parent archives) rather than direct code editing.
Multi-turn multi-modal framework for patient education combining image and text explanations with interactive dialogue.
Empirical study of token acceptance probability in tree-based speculative decoding across different task cognitive domains.
DR³-Eval benchmark for evaluating deep research agents on multimodal, multi-file report generation tasks with reproducible evaluation.
M2-PALE explains multi-agent MCTS-minimax hybrid behavior using process mining and LLMs to improve agent interpretability.
CAMO framework automates causal discovery in LLM agent simulations to understand micro-to-macro mechanisms driving emergent social outcomes.
First repository-level benchmark for evaluating LLM agents on real hardware bug repair with 417 instances from historical fixes.
LLM planning framework decoupling planning from execution via MCTS-generated atomic experience retrieval without fine-tuning.
Framework analyzing five layers of mutability in persistent language-model agents with tool use, memory, and runtime adaptation.
Essay on AI agents transforming scientific research structure, efficiency, and collaboration through information replication and sharing.
Framework coupling LLM semantic decoupling with graph contrastive learning for text-attributed graphs via orthogonal decomposition.
Self-evolving chain-of-thought generation method using genetic algorithms for data synthesis in mathematical reasoning tasks.
Benchmark for evaluating self-centric intelligence in multimodal LLMs through mirror-based embodied intelligence tasks.
Generative educational agent that simulates student cognitive evolution during learning practice with dynamic persona modeling.
Comparative analysis of CNN optimization methods for edge deployment, evaluating compression and dynamic early-exit approaches.
Studies how LLM integration in cognitive workflows affects user perception of own capabilities, introducing the LLM fallacy concept.
Redefines hallucination evaluation for medical SOAP note generation, proposing clinical abstraction and grounded inference metrics.