AI Phenomenology for Understanding Human-AI Experiences Across Eras
AI phenomenology framework examining subjective human-AI experiences beyond performance metrics and usability scales.
AI phenomenology framework examining subjective human-AI experiences beyond performance metrics and usability scales.
Research framing LLM context windows as L1 cache; proposes demand paging and virtual memory hierarchy for efficient token reuse.
Automated detection and root-cause analysis pipeline for flaky tests in quantum software systems.
PlayWorld: autonomous pipeline training action-conditioned video models for robot simulators from large-scale datasets.
WS-Net uses state-space modeling and attention mechanisms for hyperspectral image unmixing, addressing weak signal collapse in abundance estimation tasks.
Sim2Act improves simulation-to-reality transfer for robot policies using adversarial calibration and perturbation methods to handle prediction errors in decision-critical regions.
Doki is a text-native interface for generative video creation, enabling users to author videos through natural language writing instead of specialized video editing tools.
GST-VLA introduces Gaussian spatial tokenization for vision-language-action models, adding 3D geometric structure awareness to improve robot perception and decision-making.
Finetuned LLMs extract sentiment signals from textual data to forecast aluminum commodity prices, exploring when these signals are most predictive.
Survey paper on latent world models and vision-language-action systems for autonomous driving, covering taxonomy, evaluation frameworks, and challenges.
Vision-language retrieval framework for skin cancer case search using composed image-text queries with global and local representation alignment.
VIVID-Med uses frozen LLMs as structured semantic teachers for pretraining medical vision transformers, improving clinical image analysis.
Language-driven embodied navigation system using semantic priori-maps and chain-of-thought prompting for functional buildings.
Human-in-the-loop framework for post-training vision-language-action models in robotic dexterous manipulation tasks.
Research on mitigating catastrophic forgetting in class-incremental learning using causal feature expansion methods.
Rubric-guided reinforcement learning framework for dense image captioning that improves diversity and generalization over supervised distillation from VLMs.
Method learning netlist representations from LLM-generated imperfect RTL code, scaling beyond small circuits using self-correction and structural learning.
Full-duplex speech-to-speech dialogue system combining cascaded ASR-LLM-TTS without VAD segmentation, enabling natural conversational interaction.
Framework bridging discrete diffusion language models with autoregressive models to enable non-sequential global reasoning and plan revision in multi-agent systems.
Study analyzing emotion as a latent representational factor in LLM reasoning and attention mechanisms, rather than just a prediction target.
VR agent pipeline that integrates prosodic emotional context from speech into LLM dialogue processing for emotionally-aware conversational responses.
Research evaluating LLMs as interactive agents in adversarial, time-sensitive zero-sum environments, assessing strategic reasoning and decision-making beyond static benchmarks.
TaSR-RAG uses taxonomy-guided structured reasoning to improve retrieval-augmented generation systems, addressing context redundancy and multi-hop reasoning challenges in LLM-based knowledge systems.
Testing-time adaptive graph neural network for cross-domain anomaly detection addressing domain shift challenges.
Dataset condensation technique for classical clinical models enabling privacy-preserving synthetic data generation.
Contrastive learning method for skeleton-based action recognition using multi-view mini-max game framework.
Multiple instance learning approach for mammography classification using foundation model features with weak supervision.
Offline-to-online reinforcement learning method for safe robot policy alignment using action space constraints.
Competition for end-to-end document image translation combining OCR and NLP for complex layout preservation.
Fully convolutional diffusion model using ConvNets for efficient generative modeling compared to transformer alternatives.
Prompt-based document layout analysis framework using domain-specific descriptive knowledge for improved multi-domain generalization.
Benchmark evaluating gender stereotypes in LLMs across healthcare contexts with intersectional social determinants of health factors.
Motion forecasting system for autonomous vehicles handling open-world scenarios with imperfect perception and evolving object taxonomy.
Benchmark dataset and analysis of LLM bias showing models prioritize moral reasoning over commonsense knowledge.
AI agent framework for automated clinical target volume delineation in radiotherapy that adapts to guideline changes without retraining.
Vision-language-action model for autonomous driving combining perception and planning distillation to improve stability.
Normalizing flows framework for time series anomaly detection with temporal conditioning and uncertainty quantification.
Vision-language model adaptation method using evolutionary prompt learning to prevent catastrophic forgetting while maintaining parameter efficiency.
Research on portable O(1) autoregressive caching for state-space models via XLA compilation, removing hardware-specific kernel dependencies.
Research on online continual learning in transformers using routing mechanisms without catastrophic forgetting in non-stationary data streams.
Research on biologically-inspired learning algorithm addressing backpropagation limitations for complex temporal pattern recognition in cortex-like systems.
Framework for interpretable synthetic data generation using vision-language models with grounded evaluation metrics for downstream tasks.
Research on persona-adaptive prompting for evaluating multi-modal LLM agents in customer experience scenarios with dual-control interactions.
Training-free KV-Lock framework for video diffusion models improving foreground quality while maintaining background consistency.
Open-source framework for time series anomaly detection using graph neural networks with critical evaluation and standardized benchmarks.
ML research benchmarking three paradigms for automated cardiac risk classification from unstructured electronic health records using large-context LLMs.
Large-scale Vietnamese VQA dataset automatically constructed using pre-trained transformers.
Unified instruction-tuning framework for task-oriented dialog systems using schema-aware prompting.
Active learning pipeline for efficient preference data generation to improve RLHF alignment of LLMs.
Benchmark and improvement strategies for multi-audio understanding in large audio-language models.