How Do LLM Agents Think Through SQL Join Orders?
Research collaboration exploring how LLM agents optimize SQL join orders through iterative reasoning.
Research collaboration exploring how LLM agents optimize SQL join orders through iterative reasoning.
ChatHRP: AI health knowledge assistant providing answers from WHO documents with source citations.
Contral: AI-powered development environment with agent for building, debugging, and explaining code.
OpenAI deprecates GPT nano fine-tuning models, unclear migration path to GPT-5 variants.
Case study of Uber's AI budget overrun in 4 months due to misunderstanding variable vs fixed costs of AI deployment.
Benchmarks and performance data for Intel Arc Pro B70 GPU on LLM inference and video generation.
News headline about DeepSeek releasing open-source AI model sequel. Minimal detail provided.
GitRails: Open-source GitHub API proxy enabling fine-grained access control for untrusted AI agents.
xAI announces grok-voice-think-fast-1.0, voice agent API for customer support and enterprise workflows.
Anthropic tested removing Claude Code from Pro plan due to high demand; pricing adjustment to ration agentic development tool access.
arXiv paper on AI-based automated military course of action planning systems for operational decision support.
arXiv paper on evaluating content moderation AI systems using policy-grounded correctness instead of agreement metrics.
arXiv paper on co-evolving LLM decision-making and skill banks for long-horizon interactive tasks in game environments.
arXiv paper introducing diagnostics for detecting alignment faking in language models under monitoring vs unmonitored conditions.
arXiv paper on AI agents handling complex domain-specific workflows, addressing harness engineering challenges for enterprise automation tasks.
Benchmark evaluating AI agents' ability to conduct professional financial research with metrics for rigor, accuracy, and verifiability.
Adaptive test-time compute allocation framework that evolves in-context demonstrations to optimize LLM performance efficiency.
HypEHR uses hyperbolic geometry to embed EHR codes and answer clinical questions efficiently without LLM-based pipelines.
Target-based prompting method to improve demographic representation in text-to-image generative models.
Fine-tuned vision-language model for automated natural language description of embryo morphology in IVF applications.
Self-adaptive prompt engineering framework for generating task plan explanations from LLMs via systematic refinement.
Methods for measuring LLM propensity for unsanctioned behavior through environmental factor analysis and Bayesian modeling.
Analysis of AI governance compliance design and regulatory frameworks under political change. Policy-focused.
Multi-agent system for personalized physiotherapy using generative AI and computer vision for real-time pose correction.
Multi-agent empowerment framework studying emergence of complex behaviors through intrinsic motivation in agent groups.
DAVinCI framework for dual attribution and verification to reduce LLM hallucinations and improve factual accuracy.
Fine-tuning LLMs for automated online review response generation aligned with human preferences.
ReCAPA framework mitigates cascading failures in vision-language-action multi-step task systems through hierarchical correction.
Framework for auditable clinical decision support using domain-specific languages and formal verification methods.
LLM-based data augmentation method with category-aware mixture-of-experts for person-job fit in online recruitment systems.
Benchmark evaluating multimodal LLM ability to reconstruct masked text from visual context in documents and webpages without explicit prompts.
Critical analysis of MemPalace open-source memory system using spatial metaphors for LLM long-term memory with high retrieval performance claims.
Systematic evaluation of ideological bias in LLM economic causal reasoning using extended benchmark with ideology-contested intervention cases.
Reusable evaluation pipeline for AI-generated meeting summaries with public artifact package separating orchestration from task-specific evaluation logic.
Study comparing vision-language models vs LLMs on abstract visual reasoning to identify whether reasoning or representation bottlenecks limit performance.
End-to-end geocoding framework using LLMs with geohash sequence reformulation to overcome limitations of traditional multi-stage retrieval approaches.
Analysis of clock skew effects on observability and causality in distributed AI inference systems; infrastructure/observability focus.
Semantic-aware synthesis framework using specialized agent modules (analyzer, synthesizer, verifier) for improving text-to-SQL generation correctness.
Multi-agent framework addressing gender bias in machine translation quality estimation through fairness-aware agent collaboration.
Study demonstrating that brief chatbot conversations produce measurable shifts in human moral value judgments through directive AI conversations.
Multi-agent framework for long-form video understanding with hierarchical reasoning and question-aware agent collaboration for improved temporal coherence.
Live platform studying social dynamics of LLM-driven visual agents interacting in autonomous multi-agent network with image-based communication.
Framework for efficient LLM agent evaluation via diversity-guided user simulation to identify deep failure modes in multi-turn interactions.
AI agent system using reasoned tool invocation for geological lithology classification, demonstrating agentic workflow with evidence-based reasoning.
Multi-modal system extracting protein-ligand bioactivity data from scientific literature using ML on text, tables, and chemical structures.
Method using multicalibration to correct LLM prevalence estimation errors across population shifts, addressing diagnostic accuracy in varying contexts.
Action research studying team-level implementation of EU AI Act requirements in an AI startup, creating legal-text-to-action pipeline for governance.
Privacy-preserving LLM personalization architecture using composable LoRA adapters and deletable user proxies for efficient data removal without retraining.
CoFEE provides structured reasoning control for LLM-based feature discovery from unstructured data while avoiding leakage and confounding signals.
Investigates transformer generalization on symbolic reasoning tasks, showing models fail with unseen variable names and analyzing token copying difficulties.