Fashion Florence: Fine-Tuning Florence-2 for Structured Fashion Attribute Extraction
Fine-tuned Florence-2 vision-language model using LoRA for extracting structured fashion attributes from clothing images as JSON.
Fine-tuned Florence-2 vision-language model using LoRA for extracting structured fashion attributes from clothing images as JSON.
Score-based conditional energy model for inference in hybrid Bayesian networks with discrete and continuous variables.
Study of routing traces in Attention-Residual transformers for post-hoc calibration, examining whether routing signals improve uncertainty estimates.
Unified geometric framework using flag varieties to theoretically explain alignment phenomena in deep networks including gradient flow and neural collapse.
UFO framework for robust continual graph learning addressing noisy annotations in evolving graphs through flow-oriented methods.
Nautilus Compass: black-box persona drift detector for production LLM coding agents that tracks constraint violations and memory degradation over long sessions.
Method enabling efficient online learning in transformers through continuous latent contexts for multi-turn adaptation and feedback-driven decision making.
SVAR-FM framework for time series causal discovery using physics-based simulators to generate interventional data with flow matching.
EgoMemReason benchmark for evaluating memory and reasoning in ultra-long egocentric video understanding for embodied agents and smart glasses.
Key-Value Means (KVM): O(N) attention mechanism supporting fixed or growing state for long-context transformers with subquadratic prefill performance.
Identifies 'Cartesian Shortcut' vulnerability in vision reasoning benchmarks where MLLMs exploit grid-based layouts through explicit coordinate discretization.
Theoretical analysis explaining why sparse autoencoders (SAEs) show layerwise scaling variations through manifold geometry, advancing understanding of activation space structure.
VALDI benchmark demonstrating 'Pseudo-Deliberation' failure mode where LLMs show reasoning without behavioral alignment, revealing value-action gaps.
Position paper identifying 'Agentic Denominator Gaming' threat where malicious actors deploy AI agents to flood academic conferences with low-quality papers exploiting stable acceptance rates.
NaiAD dataset of 58,999 ad-embedded LLM responses with evaluation metrics for studying LLM-native advertising balancing user experience and platform revenue.
VIGOR: reward model for LLM reinforcement learning that eliminates need for external verifiers by using intrinsic gradient-norm signals, improving scalability to new tasks.
Self-training method using team-based self-play with dual adaptive weighting to improve LLM alignment while reducing dependency on human-labeled data and addressing synthetic data quality issues.
Inference-time pruning method reducing unnecessary tool calls in LLM reasoning systems while maintaining performance.
GPU-accelerated Boruta feature selection algorithms for high-dimensional datasets with improved computational efficiency.
Self-improvement framework for LLMs using intrinsic rewards without verifiers, enabling autonomous evolution on open-ended tasks.
Small 0.3B multilingual NER model for detecting 42 PII entity types across languages and document formats.
Analysis of attention drift phenomenon in speculative decoding drafters showing attention shifts from prompt to generated tokens over speculation chains.
Continual Harness framework for online adaptation of embodied agents with iterative refinement, demonstrated on Pokemon gameplay.
Specification inference tool for Move Prover combining weakest-precondition analysis with agentic Claude coding for reduced boilerplate.
Tag-based few-shot example selection method for prompting LLMs to generate causal factors and preventive measures from medical incident reports.
LLM-based framework for automated psychological crisis assessment from speech for mental health hotline support.
Vision for unified memory paradigm for agentic AI in 6G radio access networks to bridge semantic bottleneck in disaggregated architectures.
C-BPO framework for personalizing LLMs using binary preference feedback with inter-user calibration.
Guidance mechanism for stochastic interpolant robot policies enabling test-time steering without retraining for dynamic objectives.
Portable specification for multi-agent coordination as self-improving skills distributed across agent frameworks.
NCO plugin for constraining LLM decoding to prevent undesirable outputs like profanity and PII during generation without post-processing.
Metis framework reformulates LLM red teaming as policy optimization using adversarial MDPs for improved jailbreak discovery.
Test-time adaptation framework for vision-language-action robotic models that uses successful execution history to improve closed-loop reliability.
ViSRA framework for probing spatial reasoning in multi-modal LLMs through video-based agent without requiring model retraining.
Adaptive perception system for autonomous driving that dynamically allocates computation based on scene complexity.
MicroWorld framework enhances MLLMs for microscopy via multimodal attribute graphs to bridge domain gap with limited training data.
Study investigating whether scaling vision models improves localization-based explanation quality across ResNet, DenseNet, ViT architectures.
NLP method for detecting and analyzing contradictions in scientific peer reviews using fine-grained analysis beyond sentence pairs.
Security framework addressing SQL injection vulnerabilities in LLM-driven database interfaces through prompt-to-SQL translation.
MTA-RL framework combines transformers and reinforcement learning for autonomous driving with 3D scene understanding.
Neural flows method for irregular multivariate time series classification that models inter-variable interactions in one-step mapping.
Comparative study of ML vs DL for out-of-distribution detection in medical imaging, showing ML can match DL performance on constrained datasets.
LegalCiteBench evaluates citation reliability in legal LLMs, testing case authority provision to address hallucination and fabrication risks.
Protein design method combining protein language models with preference alignment to avoid catastrophic forgetting while steering toward desired functions.
Concept erasure in diffusion models via cross-attention sparsity maintaining effectiveness when scaling to larger architectures like SDXL.
Local LLM approach for automated deliberative process privilege classification in FOIA document redaction without external APIs.
Security analysis of knowledge poisoning attacks on medical multimodal RAG systems, demonstrating adversarial risks in retrieval databases.
MemReread improves agentic long-context reasoning by memory-guided rereading to recover latent evidence when processing document chunks linearly.
DP-LAC enables differentially private federated fine-tuning of LLMs with lightweight adaptive gradient clipping for privacy-preserving training.
RAG system using Qwen for Ukrainian multi-domain document understanding with PDF chunking, dense retrieval, and answer reranking.