Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
Method to enhance text embedding models' semantic reasoning through multiple forward passes, tested on benchmarks.
Method to enhance text embedding models' semantic reasoning through multiple forward passes, tested on benchmarks.
Cross-task visual in-context learning via VLMs for handling mismatched demonstrations. Extends VICL to tasks differing from examples.
Benchmark and evaluation of foresight intelligence in vision-language models for anticipating future events. New VQA dataset for predictive capabilities.
Federated learning framework for multimodal feature extraction with contrastive learning under non-IID data. ML research with privacy focus.
Hierarchical Mixture-of-Experts framework for vision-language-action policies handling heterogeneous robot data. Enables generalist multimodal policies.
Compression technique for Mixture of Experts models using structured butterfly matrices. Reduces memory scaling for efficient edge deployment.
Token compression technique for omnimodal LLMs using audio-driven semantic chunking. Improves inference efficiency for multimodal models.
Framework for asynchronous multi-agent collaboration on long-horizon software engineering tasks. Addresses agent coordination and timely completion.
LLM-informed planning framework for object search in partially-known environments using prompt selection. Combines planning with LLM knowledge.
Multimodal LLM framework for annotating broadcast television content. Domain-specific application with limited general relevance.
Study of LLM reasoning modes in multi-agent negotiation simulations. Examines behavior reproduction vs optimal solving in agent interactions.
LLM-based agent framework that generates proof-of-concept tests to validate bug detection reports. Combines agents with automated testing.
Self-evolving memory system for LLM-based code generation on private libraries using execution feedback. Improves code generation with enterprise context.
Semantic search system for clinical notes at health system scale using embeddings. LLM application but domain-specific healthcare focus.
Benchmark for LLM memory retrieval precision showing current evaluations mask severe precision failures through complete belief dumps.
Survey of mathematical reasoning in LLMs covering benchmarks, architectures, evaluation methods, and open challenges in the field.
Pass@K optimization research for code generation improving test-time compute allocation by coordinating diverse sampling instead of independent draws.
Research on scaling reinforcement learning from verifiable rewards for agentic LLMs using synthetic task augmentation instead of human curation.
Research showing LLMs fail to verify source quality during multi-source synthesis despite detecting fabrication in isolation.
TLA-Prover: 20B parameter LLM trained via preference optimization to generate formally verifiable TLA+ specifications for distributed systems.
MetaConfigurator extends JSON Schema editor with RDF authoring for semantic interoperability in scientific workflow data.
Training-free concept detection and steering in transformer models by analyzing sign patterns in raw transformer dimensions without learned dictionaries.
Study analyzing variability loss in AI-generated code from LLM vibe coding, showing programs have minimal compile/runtime variability.
Polycepta improves multi-object tracking with dynamic object-centric appearance estimation to complement motion prediction.
RWGBench introduces benchmark for evaluating related work generation in academic papers using citation-level scholarly positioning metrics.
Empirical study of ERC-8004 decentralized AI agent protocol, analyzing trust mechanisms in permissionless agent economies.
JuZhou 1.0 is ultra-lightweight text-to-image model designed for edge deployment and offline execution on China-developed AI accelerators.
MultAttnAttrib provides training-free multimodal attribution for long-document QA, improving interpretability and safety in AI assistants.
RoboDojo provides unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks.
Wan-Streamer v0.2 improves resolution of native-streaming audio-visual interaction model while maintaining 200ms latency.
SearchGen uses agentic visual generation with web search to overcome knowledge cutoff limitations in text-to-image models for open-ended user requests.
Weak-to-Strong Generalization enables cheaper RL post-training by running RL on small models then distilling to larger models for improved reasoning.
Floor-First Triage proposes analytical estimation methods for LLM serving optimization to replace grid-search approaches and improve latency profiling workflows.
UBEP optimizes Mixture-of-Experts model communication on high-bandwidth superpods by addressing execution serialization and bandwidth bottlenecks in production deployments.
TriRoute jointly optimizes attention resolution, expert selection, and KV-cache allocation in language models using learned routing to decouple model quality from per-token inference cost.
Open-source foundation model for wearable motion sensing with comprehensive study of pretraining and scaling principles.
LLM-guided time-series forecasting approach leveraging process documentation for industrial soft sensing with scarce labeled data.
Method for adapting specialist industrial models to new scenarios using LLM-guided reasoning without parameter modification.
Mechanistic study of reward valuation in vision-language models linked to anhedonia assessment from clinical psychology.
Approximation ratio analysis for greedy algorithm in myopic Bayesian active learning for linear regression with tight bounds.
DsrFGW: Optimal transport method for graph matching combining node features and structure via diffusion-inspired approach for sparse/noisy graphs.
Analysis of latent reasoning faithfulness in hidden state reasoning across training trajectories, showing unfaithful behaviors beyond converged checkpoints.
MESH-FL: Entropy-guided tensor compression for multimodal federated learning on edge devices accounting for modality-specific spectral differences.
Generative method for temporal point processes using rough path signatures as feature maps, addressing sequence-level evaluation limitations.
FedDualAtt: Personalized federated learning for ECG classification using split transformer attention heads with global and local branches.
Study on human-AI complementarity under asymmetric information, analyzing when human decision makers fail to realize gains from ML model augmentation.
Meta-learning approach for learned optimizers that efficiently scales to long-horizon inner problems, improving upon hand-designed optimizers like Adam.
Bayesian deep ensemble method for predictive regression combining statistical rigor with scalability and calibrated uncertainty estimates.
Knowledge distillation for time series classification, transferring knowledge from large teacher to efficient student model for resource-limited environments.
Analysis of signals predicting correctness in text-to-SQL generation using self-consistency and schema-relevance metrics on BIRD and Spider benchmarks.