Caliper-in-the-Loop: Black-Box Optimization for Hyperledger Fabric Performance Tuning
Bayesian optimization with dimensionality reduction for automated Hyperledger Fabric performance tuning via black-box benchmarking.
Bayesian optimization with dimensionality reduction for automated Hyperledger Fabric performance tuning via black-box benchmarking.
Qwen3-32B finetuning for multilingual polarization detection across detection, type, and manifestation subtasks.
PIEGraph combines graph neural networks with analytical physics for data-efficient deformable object dynamics learning in robotics.
AI tutor system treating pair programming collaboration as instructional target using joint visual attention and cognitive modeling.
Qwen3-32B finetuning with data augmentation and self-training for Reddit conspiracy detection classification task.
Vision-language model for grounded reasoning using perceptual flow networks to reduce hallucination and language bias.
Systematic audit of technical debt and code quality in LLM and agent-generated software, revealing machine-specific defect patterns.
Multimodal LLM grounding molecular reasoning in chemical structures via Morgan fingerprint embeddings for drug discovery auditing.
Compact one-step diffusion model for efficient image super-resolution via task-optimal backbone discovery.
Diffusion-based planner for offline safe RL with decoupled guidance to balance reward and constraint satisfaction adaptively.
Pattern-based AI-assisted methodology for sensor-driven application development using Pegasus workflows on FABRIC testbed.
Methodology for rapid sensor data application development across edge-to-cloud infrastructure using pattern-based approaches.
Uses SHAP analysis to decompose RL algorithm and hyperparameter contributions to generalization gaps in robotics, improving configuration selection.
SpecKV adaptive speculative decoding with compression-aware gamma selection for LLM inference acceleration. Research on optimizing draft model token proposal length.
Algorithm for logic-constrained shortest path problem with satisfiability constraints applied to flight planning optimization. Specialized constraint solving research.
Survey of LLM-powered AI agent systems and industrial applications with multimodal capabilities. Comprehensive research on agent architecture and deployment.
MSEarth benchmark dataset for multimodal LLMs on earth science phenomenon discovery. Research-backed evaluation framework for MLLM scientific reasoning.
Theoretical analysis of neural networks as iterated belief change systems using Darwiche-Pearl framework.
NaviGNN hybrid framework combines agent-based modeling, RL, and GNNs for multi-agent navigation in futuristic smart city topologies.
SMoE algorithm-system co-design deploys Mixture of Experts models on edge devices via expert importance-guided substitution.
Chain-of-Thought Reconstruction technique prunes reasoning LLMs while maintaining performance by reconstructing thinking tokens.
EngiBench hierarchical benchmark evaluates LLMs on real-world engineering problem-solving with uncertainty and open-ended settings.
oMeBench benchmark evaluates LLM reasoning capabilities on organic chemistry mechanism elucidation tasks.
How² method enables agents to learn from procedural how-to questions to reduce uncertainty and fill knowledge gaps in planning.
Geospatial Awareness Layer grounds LLM agents in structured earth data for wildfire response planning and disaster management.
GOAT training framework for LLM agents with tool use, synthesizes API execution data from documentation for fine-tuning without human annotation.
Reject option prediction method quantifying both epistemic and aleatoric uncertainty for high-stakes prediction applications.
Agentic learner with multimodal semantic memory that grows and refines domain knowledge across multiple modalities to avoid repeating mistakes.
SimVC-CAS collective agent system simulates venture capital decisions through multi-agent roleplay to predict startup success.
SCALER framework uses synthetic adaptive learning environments with RL to enhance LLM reasoning by maintaining alignment between task difficulty and model capability.
Framework characterizing homogenization in LLMs through bias amplification and mode collapse, proposing metrics for AI safety evaluation.
Multi-agent pipeline using LLMs to convert customer reviews into actionable business recommendations through hierarchical decision-support.
MemeLens is a multilingual multitask Vision Language Model for meme understanding across hate detection, propaganda, and sentiment tasks.
Multi-Agent Actor Critic framework enables decentralized LLM collaboration with parallel inference and flexible deployments without predefined execution protocols.
Dynamic One-Shot Policy Refinement method reduces resource costs for reinforcement learning of LLM reasoning while maintaining performance on complex tasks.
MINT framework for AI agents to actively elicit human inputs through optimal interaction strategies in joint planning with incomplete information.
Introduces Agentic Workflow Reconstruction (AWR) task to synthesize explicit, interpretable workflows from opaque LLM-based agentic systems.
Analyzes 809 LLMs released 2022-2025 via scaling-law regressions, finding developer-specific efficiency advantages beyond raw compute scaling.
Proposes ABD benchmark for default-exception abduction in finite first-order worlds with SMT verification across three observation regimes.
Introduces VeRO, an evaluation harness for coding agents optimizing other agents through iterative edit-execute-evaluate cycles.
Presents agentic multi-source grounding approach for disambiguating user queries in multi-category marketplaces, tested at DoorDash.
Introduces TRACED framework evaluating LLM reasoning quality through geometric kinematics measuring progress and stability of reasoning traces.
Establishes quantitative theory of grokking as norm-driven representational phase transition, deriving the Norm-Separation Delay Law.
Augments LLMs with computational argumentation for explainable, contestable decision-making in high-stakes domains.
Reveals safety degradation in large reasoning models occurs after chain-of-thought generation and proposes safety decisions before CoT.
Proposes SpecTM, physics-informed masking for Earth observation foundation models enforcing spectral constraints during reconstruction.
Introduces WMF-AM, a probe measuring LLM working memory via cumulative state tracking across sequential operations without scratchpad.
Develops category-theoretic framework for defining, comparing, and analyzing AGI systems and benchmarks.
Introduces ScoringBench, an open benchmark evaluating tabular foundation models using proper scoring rules instead of point-estimate metrics.
Presents neurosymbolic architecture using ontology-constrained reasoning in Foundation AgenticOS to reduce hallucination and enforce compliance in enterprise LLM agents.