Towards Context-Aware Image Anonymization with Multi-Agent Reasoning
CAIAMAR: multi-agent framework for context-aware image anonymization in street-level imagery using agentic reasoning.
CAIAMAR: multi-agent framework for context-aware image anonymization in street-level imagery using agentic reasoning.
Kill-chain canary methodology for tracking prompt injection attacks across multi-agent LLM systems with stage-level diagnostics.
System for making mathematical theorems interactive by grounding LLM-generated explanations in formal representations enabling execution and stepping.
Framework for eliciting and verbalizing LLM assumptions to explain and mitigate sycophantic behavior in model outputs.
Multi-stage LLM-assisted workflow for scientific algorithm development separating theory extraction, formal specification, and code generation.
Method for LLM personalization using a small portfolio of models capturing diverse user preferences without per-user models.
Distributional reinforcement learning approach for decision-making in healthcare, accounting for uncertainty across heterogeneous populations.
ALTO: system for adaptive LoRA hyperparameter tuning and orchestration across heterogeneous LLM fine-tuning workloads in multi-tenant environments.
WisdomInterrogatory (LuWen): open-source Chinese legal language model built on Baichuan foundation model for legal domain applications.
System for safe capability evolution in embodied agents with compatibility checking and runtime rollback mechanisms.
Training-free open-vocabulary semantic segmentation framework (OV-Stitcher) leveraging pretrained vision-language models without additional training.
HyperMem: hypergraph-based memory architecture for conversational agents enabling long-term context tracking and high-order associations.
Physics-aligned simulator (SIM1) for generating synthetic data in deformable object robotic manipulation tasks.
Framework combining LLMs with graph neural networks for text-attributed graph learning in low-resource settings using GNN feedback.
QuanBench+ unified benchmark for LLM quantum code generation across Qiskit, PennyLane, Cirq with 42 aligned executable tasks.
Benchmark evaluating robustness of LLM reasoning with 14 perturbation techniques applied to mathematical reasoning tasks.
DRTO combines token-level RLHF with distributional robustness to improve LLM resilience to input perturbations and formatting changes.
Automated label function generation for data annotation using LLMs with structured exploration-exploitation strategy.
TinyML Z-score anomaly detection system running on resource-constrained microcontrollers using power side-channel data.
CSAttention: sparse attention mechanism for accelerating LLM inference by reducing KV-cache bottlenecks through centroid-scoring without retraining.
Framework evaluating when LLMs should act versus escalate decisions using uncertainty estimation across five real-world domains.
AlphaLab autonomous research system using frontier LLMs as agents to automate full experimental cycles in optimization domains without human intervention.
StructRL: recovers dynamic programming structure from distributional RL learning dynamics. Bridges data-driven and structured approaches for stable learning.
Bayesian inference for spiking neural networks in speech processing. Explores weight uncertainty and loss landscape smoothing for temporal tasks.
Evidential Transformation Network: adapts pretrained models for post-hoc uncertainty estimation. Efficient alternative to ensembles/MC dropout for deployed models.
VOLTA: benchmark comparing uncertainty quantification methods for deep learning. Evaluates 10 UQ baselines across modalities and distribution shifts.
Game-theoretic analysis of creator incentives in multi-agent recommender systems. Cooperative game formulation for fair collaboration in bandit problems.
PRAGMA: foundation models for banking event sequences. Transformer-based architecture with self-supervised pretraining on financial transaction data.
Skip-Connected Policy Optimization (SKPO) for reinforcement learning with reasoning tasks. Improves upon GRPO by addressing high-variance advantage estimation.
EvoLen: evolution-guided tokenization approach for DNA language models. Addresses fundamental tokenization design challenges in biological sequence modeling.
Experience replay for LLM post-training RL formalizing optimal buffer design as trade-off between sample efficiency and data freshness.
Tensor decomposition method quantifying uncertainty in LLM-based multi-agent systems accounting for communication and role dependencies.
CLOVER framework for multi-agent RL cooperation conditioning value decomposition on realistic wireless communication graphs.
LottaLoRA training paradigm showing frozen random backbones with trained LoRA adapters recover 96-100% performance across diverse tasks.
Adaptive simulation experiment framework using pairwise comparisons to optimize LLM policies for operations management tasks.
Prompt optimization method decomposing reward variance into response and prompt variance to identify task amenability to optimization.
Multi-agent actor-critic reinforcement learning for disaster resilience controlling power, communication, and emergency response systems.
4-bit floating-point format (HiFloat4) for efficient language model pre-training on Ascend NPU hardware.
Guidance method for consistency models using joint flow distribution learning to enable classifier-free guidance without separate teacher model.
Training curriculum method for discrete flow-based image generation models to improve one-step sampling stability and quality.
Analysis of LoRA adapter spectral geometry to identify fine-tuning objectives and predict harmful model behavior in language models.
Safety steering mechanism for multimodal LLMs using dictionary-aligned concept control to prevent unsafe outputs without retraining.
Theoretical analysis of finite-sample properties and identifiability bounds for nonlinear Independent Component Analysis algorithms.
Demonstrates power-law scaling of classification error with number of classes and how chain-of-thought decomposition reduces error through task splitting.
Practical analysis of chain-of-thought distillation from students to teachers, revisiting capacity gap assumptions and baseline comparisons.
Conformal prediction framework for transformers providing uncertainty quantification and calibration for trustworthy LLM deployment.
Analysis of causal inference applications in graph representation learning and risks of aggregating graph elements.
Adaptive Thompson sampling for high-dimensional Bayesian optimization addressing sparse candidate point grids.
Dynamic policy optimization bridging SFT and RL for LLMs, addressing bias-variance tradeoff in post-training through adaptive loss weighting.
Empirical study on effectiveness of advanced optimizers for multi-task learning, identifying overlooked factors in optimization approaches.