UI-Venus-1.5 Technical Report
UI-Venus-1.5: GUI agent with 2B, 8B, and 30B-A3B variants for automating digital environment interactions with broad generality and strong task performance.
UI-Venus-1.5: GUI agent with 2B, 8B, and 30B-A3B variants for automating digital environment interactions with broad generality and strong task performance.
VESPO: reinforcement learning method for training LLMs with improved stability through soft policy optimization and importance sampling to address policy divergence.
KBVQ-MoE: compression technique for Mixture of Experts LLMs using vector quantization and SVD to reduce parameter size and memory for resource-constrained deployment.
Sim2Radar uses VLM-guided scene reconstruction to synthesize radar training data from RGB images, addressing radar dataset scarcity.
Identifies and analyzes 'silent inconsistency' problem in data-parallel LLM fine-tuning where gradient synchronization doesn't ensure worker-level optimization alignment.
ST-EVO framework for self-evolving LLM-powered multi-agent systems that dynamically construct task-adaptive communication topologies instead of predefined structures.
Studies mechanical basis of capability emergence in neural networks across scales 405K-85M parameters, finding scale-invariant representation collapse and top-down reorganization.
Proposes AI-CARE metric incorporating carbon emissions and energy consumption alongside standard performance metrics for ML model evaluation.
Analyzes quality issues in AI safety datasets, finding they rely on superficial 'triggering cues' rather than genuine adversarial patterns.
Randomized trial showing AI-generated feedback suggestions via FeedbackWriter improve student revisions when reviewed by human TAs in economics course.
Proposes symbolic alternative to GNN message-passing for more interpretable and expressive graph learning in high-stakes domains.
MASPO algorithm improves LLM reasoning through reinforcement learning with verifiable rewards, addressing gradient utilization and probability mass issues in existing RLVR methods.
User study comparing chatbots, games, and essays for persuasive learning on sustainability topics with identical content.
Uses LLM-assisted reasoning to map 2D engineering drawing annotations to 3D CAD features for manufacturing automation and process planning.
Benchmark study (SP-ABCBench) evaluating whether LLM agents can simulate human security and privacy attitudes and behaviors for risk forecasting.
Examines integration of AI into science education materials, covering personalization, adaptive instruction, and accessibility in K-12 learning contexts.
Proposes Reasoning Processing Unit (RPU) architecture to address memory bandwidth bottlenecks in LLM inference, particularly for reasoning applications with long outputs.
Study revealing performance degradation when using state-of-the-art text-to-image models as synthetic training data generators for vision tasks.
Benchmark suite for evaluating video reasoning capabilities in modern video models including spatiotemporal reasoning and scaling behavior.
MoBiQuant enables elastic LLM deployment with token-adaptive mixture-of-bits quantization supporting dynamic precision switching at runtime.
Hybrid-policy reinforcement learning framework for multi-modal LLMs with exploration strategy to prevent entropy collapse during RL training.
Framework addressing class imbalance, overlap, and noise in multi-class learning using regional partitioning and meta-heuristic ensembles.
Layer gradient analysis method for identifying optimal layers for knowledge editing in LLMs while preserving model behavior.
ESM framework for merging multiple task-specific fine-tuned models using principal component analysis to reduce task interference.
KnapSpec framework reformulates self-speculative decoding as knapsack problem to optimize LLM inference throughput through adaptive layer selection.
Extension of TabPFN foundation model to handle multimodal tabular data integrating images, text, and tables in unified framework.
Convex optimization-based clustering algorithm with LLM integration for analyzing biomedical literature and detecting trends in anti-aging research.
Study of LLM truthfulness representations across domain-general and domain-specific directions using probe generalization across five truth types.
Discrete diffusion framework using sample-efficient estimators for generative modeling over discrete state spaces with conditional probabilities.
Curriculum learning approach that recursively decomposes complex datasets into simpler components using teacher-student framework with step-by-step reasoning.
Time-series foundation models augmented with in-context learning to adapt to unseen tasks without fine-tuning.
QuantVLA: post-training quantization framework for Vision-Language-Action models to reduce compute and memory demands for embodied AI agents.
CaDrift: synthetic data generator using Structural Causal Models to create data streams with controlled distributional and covariate shifts for evaluating ML methods.
Analysis of representation geometry dynamics during chain-of-thought reasoning in LLMs using manifold capacity theory.
Self-supervised graph learning method incorporating fragment-level information for molecular representation learning.
Plug-and-play guidance method for flow-based generative models improving sample fidelity without doubling inference cost.
Method incorporating causal context into Shapley values for accurate multivariate feature importance measurement.
Pre-training framework leveraging geometric data for efficient neural physics simulation with better transfer.
Analysis of safety challenges in unsupervised elicitation techniques for steering language models toward truthful outputs.
Online learning algorithm with Wasserstein distributionally robust optimization for risk-averse sequential decisions.
Framework for active exploration and model estimation in tabular MDPs based on coverage-based complexity.
Certification method for verifying DNN ownership against model extraction attacks in MLaaS systems.
Differentiable framework using Gaussian reparameterization for scheduling optimization in compilation and synthesis.
Analysis of how protein language models diverge from natural language transformers with improved inference methods.
Online alignment method for LLMs under misspecified preference feedback, extending SAIL framework.
Attention Neural Teaching paradigm to reduce training costs for transformer-based attention learners.
Neural network pruning technique balancing compression and information preservation in fully-connected networks.
Research on normalizing flows and invertible neural networks for generative modeling and inverse problems.
Decentralized federated learning approach for multi-task LLM fine-tuning using sparse-orthogonal LoRA over wireless connections.
Apprenticeship learning framework for intelligent tutoring systems addressing sample efficiency and reward function design in educational RL.