Automated Microservice Pattern Instance Detection Using Infrastructure-as-Code Artifacts and Large Language Models
Using LLMs to detect microservice architecture patterns from Infrastructure-as-Code artifacts for documentation purposes.
Using LLMs to detect microservice architecture patterns from Infrastructure-as-Code artifacts for documentation purposes.
Analysis of 1.8M Hugging Face models tracking how multimodal capabilities emerge and propagate across open LLM families.
Systematic evaluation of four prompting strategies across GPT models on chart-based question answering, isolating prompt structure effects.
MERIT system combining memory and retrieval mechanisms with LLMs for interpretable knowledge tracing in educational settings.
TIPS framework for training search-augmented LLMs with reinforcement learning, improving credit assignment and reward shaping for question answering tasks.
Methods for generating high-quality synthetic training data using LLMs to fine-tune smaller models, analyzing diversity and distribution in embedding space.
Mechanistic interpretability study investigating whether LLMs develop genuine emotional representations or merely detect emotion keywords through circuit analysis.
Compact uncertainty estimation method for LLMs scoring cross-layer agreement patterns in internal representations via single forward pass.
Sparse Feature Attention method reducing transformer self-attention cost via k-sparse feature representations instead of sequence-level sparsity.
Mathematical framework interpreting LLM hidden states as points on latent semantic manifolds with Riemannian geometry and Voronoi partitions.
Training-free hallucination detector for LLMs using sample transform cost to measure output distribution complexity without fine-tuning.
Chinese financial news dataset and benchmark for evaluating LLM-based agents in macro and sector asset allocation decision-making.
Decision Transformer approach for optimizing emergency vehicle signal preemption using offline, return-conditioned sequence modeling.
Geometric Mixture-of-Experts framework for graph representation learning using curvature-guided routing on heterogeneous topologies.
Dataset aligning instruction manuals with assembly videos for evaluating multimodal LLMs on real-world technical tasks.
AEGIS infrastructure for governance of adaptive medical AI systems under FDA and EU regulations with continuous improvement.
Multi-task deep learning framework for predicting lithium-ion battery state-of-health and remaining useful life.
Delta-Aware Quantization framework for post-training LLM compression that preserves knowledge from alignment fine-tuning.
Classification approach for wind power ramp event forecasting under severe class imbalance for grid stability.
AgentSLR uses agentic AI to automate systematic literature reviews in epidemiology from retrieval through synthesis.
Method for adding trained persistent memory to frozen decoder-only LLMs without cross-attention mechanisms.
Applies conformal prediction for formal safety guarantees in wildfire spread prediction using tabular, spatial, and graph models.
Comprehensive study of LLM-based data imputation across multiple models and datasets, analyzing hallucination effects and control mechanisms.
Combines graph signal processing with Mamba2 state-space models to create adaptive filter banks for language modeling.
Causal Direct Preference Optimization method for training LLMs to generate recommendations while mitigating spurious correlations.
Graph RAG framework combining labeled property graphs and RDF for retrieval-augmented generation over structured and semi-structured data.
T-MAP uses evolutionary search to red-team LLM agents by exploiting multi-step tool execution vulnerabilities in MCP ecosystems.
Analysis of feature importance bias in gradient boosting models under multicollinearity, affecting SHAP-based explanations.
WIST framework uses web-grounded iterative self-play with reinforcement learning to improve LLM reasoning in specific domains.
Study on using LLMs for algorithm synthesis with provable guarantees, combining mathematical reasoning with practical performance.
Research on improving conditional modeling in diffusion models, establishing equivalence between classifier-free guidance and alignment objectives.
Proposes Reasoner-Executor-Synthesizer architecture for LLM agents that maintains O(1) context window while avoiding hallucination and token cost scaling.
Research evaluating Vision-Language Models' ability to detect misleading data visualizations and deceptive captions in charts.
FAAR quantization method for NVFP4 ultra-low-bit format that adapts rounding to non-uniform numerical grid for efficient LLM edge deployment.
Study on multimodal fusion strategies for time series forecasting showing naive fusion fails and proposing constrained fusion approach.
MTEO method for few-step diffusion sampling by distilling layer-wise, step-wise time embeddings to accelerate inference.
AI Co-Scientist framework combining LLM agents with cloud computing to automate search ranking research from ideation through GPU training.
Cross-task evaluation study of LoRA adapters showing nominal instruction-tuning labels don't reliably predict realized instruction-following capabilities.
Symbolic Graph Network framework for discovering partial differential equations from noisy sparse data without numerical differentiation.
Adaptive temporal control system for autonomous agents that learns optimal action intervals using hyperbolic geometry predictive signals.
Open-source framework (CaP-X) for benchmarking and improving code-as-policy agents for robot manipulation tasks.
Token-level analysis of distributional shifts in RLVR fine-tuning of LLMs to understand mechanisms underlying reasoning improvements.
LLM-guided headline rewriting system that enhances reader engagement while maintaining editorial integrity and avoiding clickbait.
Ablation study analyzing specialization patterns in hybrid language models combining attention with state space models on sub-1B parameter models.
Framework for building language model general capabilities via automatic curriculum of cross-entropy game tasks for relevant skill discovery.
Inference-time scaling method using small latent verifiers instead of multimodal LLMs to score and select outputs while reducing computational cost.
Empirical study measuring semantic novelty of 13,847 IS papers (2020-2025) to assess whether LLM productivity gains translate to genuine intellectual advancement.
LLMON proposes a markup language for LLMs that preserves structure and semantics in prompts, distinguishing between instructions and data in input/output.
ChatP&ID is an agentic RAG framework enabling LLM interaction with engineering P&ID diagrams using knowledge graphs for cost-effective grounded reasoning.
Ego2Web benchmark for multimodal web agents grounded in egocentric video, evaluating agents performing real-world workflows with physical context awareness.