Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
MeasureBench benchmarks vision-language models on visual measurement reading tasks with real-world and synthesized instrument images.
MeasureBench benchmarks vision-language models on visual measurement reading tasks with real-world and synthesized instrument images.
Xmera framework evaluates adversarial man-in-the-middle attacks on LLM factual recall through prompt injection, measuring vulnerability of question-answering systems.
MOON2.0 addresses multimodal imbalance in MLLMs for e-commerce product understanding through dynamic modality-balanced representation learning.
HumorChain: Theory-guided multi-stage reasoning framework for interpretable multimodal humor generation using LLMs.
Study on spatial reasoning in LLMs for 3D scene understanding, examining attention masking mechanisms for order-agnostic objects.
ThinkDeeper: Framework for autonomous vehicle grounding using world models for 3D spatial reasoning and scene prediction.
Research on metaphor-based jailbreak attacks against text-to-image models' safety defense mechanisms.
Zero-shot object navigation for robots using ensemble prediction of future states in unseen, cluttered environments.
Empirical study examining reproducibility gaps in code generated by LLM coding agents and missing dependency specifications.
VLM-CAD: Collaborative agent design workflow for analog circuit sizing using vision-language models with spatial reasoning.
Information-theoretic analysis of trade-offs between fairness, privacy, and accuracy in machine learning using Chernoff Information.
HAVEN: Framework for long-video understanding using agentic search and audiovisual entity cohesion to maintain global coherence.
Analysis of representational homomorphism in transformers to predict and improve compositional generalization in language models.
Vision-DeepResearch: Framework augmenting multimodal LLMs with tool-calling capabilities for visual and textual search.
1S-DAug: One-shot data augmentation method for improved few-shot learning generalization using generative synthesis.
Residual Decoding: Training method to reduce hallucinations in vision-language models using history-aware residual guidance.
FlyPrompt: Brain-inspired routing method for continual learning from non-stationary data streams without task boundaries.
Study evaluating behavioral consistency of LLM agents in stock market simulations against real market participant behavior.
Energy-aware reinforcement learning for robotic manipulation of articulated objects in infrastructure maintenance and smart cities.
KDFlow: Framework for efficient knowledge distillation of large language models into smaller models with heterogeneous training backends.
MA-RAG: Multi-round agentic RAG system for medical reasoning with LLMs, addressing hallucinations and outdated knowledge through iterative refinement.
Augmenting Proximal Policy Optimization with temporal sequence models for robust reinforcement learning under sensor drift and partial observability.
NCCL EP, unified communication API for mixture-of-experts architectures in large language models built on NCCL with GPU-initiated RDMA.
Training-free fine-grained visual recognition using large vision-language models with sample-wise adaptive reasoning for subordinate-level category disambiguation.
Video world models for robotics using inverse dynamics rewards to align generated trajectories with executable robot actions.
Systematic analysis of Elastic Weight Consolidation for continual learning showing suboptimal performance and proposing improvements to weight importance estimation.
Benchmark comparing PETNN, KAN, and classical deep learning models on Burmese handwritten digit recognition dataset.
Mi:dm K 2.5 Pro, 32B parameter enterprise LLM supporting multi-step reasoning, long-context understanding, and agentic workflows in Korean and domain-specific applications.
Formal specification for admission control governance of autonomous agents in institutional B2B environments with cryptographic validation.
Graph-based memory system for LLM reward prediction requiring limited labeled data for reinforcement learning post-training.
Research on chain-of-thought faithfulness in LLMs showing measurement methodology significantly affects reported faithfulness scores across 12 open-weight models.
KidGym: 2D grid-based reasoning benchmark evaluating MLLMs on spatial intelligence inspired by Wechsler Intelligence Scales.
CRoCoDiL: Continuous semantic space diffusion model for non-autoregressive language generation with improved coherence.
Industrial-scale RAG framework evaluated on automotive manufacturing requirements engineering with unstructured heterogeneous documentation.
Memory-Keyed Attention: Efficient attention mechanism reducing KV cache memory for long-context LLM inference and training.
TRACE: Multi-agent system using autonomous reasoning for seismological analysis of earthquake mechanisms from geophysical observations.
Unsupervised self-evolution training framework for multimodal LLMs achieving reasoning improvements without annotated data.
DeepXplain: Explainable deep reinforcement learning framework for multi-stage APT cyber defense with provenance graphs.
Case study on LLM-powered workflow optimization for multidisciplinary software development in automotive industry.
mSFT: Iterative algorithm addressing overfitting in multi-task supervised fine-tuning by heterogeneous data mixture optimization.
arXiv paper on safe offline reinforcement learning with budget constraints. Addresses safety-reward trade-offs in sequential decision making.
Research on synthetic data generation using LLMs to improve smaller model fine-tuning. Analyzes diversity and distribution in embedding space.
Uncertainty estimation method for LLMs using intra-layer local information scores from cross-layer agreement patterns.
Sparse Feature Attention method reducing transformer self-attention complexity through feature-level sparsity instead of sequence-level sparsity.
Mathematical framework interpreting LLM hidden states as points on latent semantic manifolds with Riemannian geometry.
Training-free hallucination detector for LLMs using sample transform cost to measure output distribution complexity.
Progressive Quantization method for robust vector tokenization in multimodal LLMs and diffusion models.
Chinese financial news dataset and benchmark for evaluating LLMs as autonomous agents in macro and sector asset allocation.
UniFluids: conditional flow-matching framework using diffusion Transformers to unify learning solution operators across diverse PDEs.
Decision Transformer approach for offline emergency vehicle signal preemption optimization without online exploration.