One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
Novel video tokenization approach using panoptic sub-object trajectories to reduce tokens for transformer-based long video processing.
Novel video tokenization approach using panoptic sub-object trajectories to reduce tokens for transformer-based long video processing.
arXiv paper on federated learning fairness evaluation, addressing discrimination at client level despite global model fairness.
Causal discovery method for irregularly sampled time series with consistency guarantees, addressing missing data challenges.
Q-chunking algorithm for improving reinforcement learning on long-horizon sparse-reward tasks in offline-to-online settings.
TinyTroupe toolkit for LLM-powered multiagent persona simulation with fine-grained specifications for realistic human behavior simulation.
Experimental study of multi-agent LLM systems for qualitative coding, examining how persona and temperature affect consensus-building vs single agents.
Argues LLM evaluations on human-designed tests are ontologically flawed; proposes AI-specific evaluation frameworks instead of anthropomorphizing model performance.
MECAT benchmark for fine-grained audio understanding using multi-expert annotations to distinguish detailed model outputs.
EventTSF integrates natural language events with time series data for improved non-stationary forecasting.
Subjective logic framework for assessing trustworthiness of AI training datasets regarding bias and fairness.
Multimodal representation learning conditioned on semantic relations beyond single shared embeddings like CLIP.
GUARD testing framework translating ethics guidelines into actionable jailbreak diagnostics for LLM safety evaluation.
Scam2Prompt framework auditing malicious endpoints in production LLMs for security vulnerabilities from training data.
Top-H decoding method balancing creativity and coherence in LLM text generation via bounded entropy.
AU-Harness open-source toolkit for standardized evaluation of audio language models with multi-turn dialogue support.
Analysis of why SFT+RL post-training works: SFT peaks on OOD early then declines; RL recovers generalization.
Token-Aware Phase Attention (TAPA) positional encoding improving long-context modeling beyond RoPE limitations.
Elastic MoE method enabling runtime scaling of activated experts for heterogeneous hardware and varying workloads.
Benchmark and mitigation techniques for sycophancy in medical vision language models for visual QA tasks.
GLAI architectural block separating structural and quantitative knowledge for accelerated neural network training.
Security analysis of information leakage from LaTeX sources and metadata in arXiv preprints using LLMs.
Multi-source reasoning alignment for MLLMs addressing concept drift in non-stationary environments via constraint satisfaction.
PAINET transformer for 3D dynamics modeling in multi-body systems using geometric symmetries for trajectory prediction.
EverydayMMQA dataset for multilingual/multimodal visual QA with cultural grounding using OASIS framework for low-resource languages.
Framework augmenting auto-regressive diffusion models with offline-trained controllers for guided generation via stepwise corrections.
Test-time scaling in LLMs reduces safety and reliability when candidate diversity decreases, revealing failure mode.
Learning from constrained demonstrations for robots with limited control interfaces via policy learning.
Batch Bayesian active learning with partial label sampling addressing scalability of acquisition functions.
User-in-the-loop pairwise preference collection for LLM alignment training without professional annotators.
GenCellAgent multi-agent framework for training-free cellular image segmentation using vision-language models with planner-executor-evaluator loop.
LILO framework uses LLMs to translate natural language feedback into structured preference signals for Bayesian optimization.
ActivationReasoning framework enables systematic logical reasoning in LLM latent spaces using sparse autoencoders for interpretability.
SAID defense mechanism against jailbreaks via prefix probing for safety-aligned LLMs without external filtering.
BenchPress evaluation suite and strong simple baselines for context compression in retrieval-augmented generation systems.
TetraJet-v2 enables 4-bit fully-quantized LLM training using NVFP4 format with oscillation suppression and outlier control.
Semantic Information Theory for LLMs replacing classical information theory bits with tokens, developing first-principles theory.
Continuum improves multi-turn LLM agent inference efficiency via KV cache time-to-live management across tool-interleaved calls.
COGNOS addresses MSE loss limitations in time series anomaly detection via constrained Gaussian-noise optimization and smoothing.
MURPHY improves GRPO training with feedback-aware retrospective credit assignment for multi-turn agentic code generation tasks.
Comparative analysis of Pandas, Polars, and Dask dataframe libraries in deep learning pipelines focusing on energy consumption and GPU interaction.
Membership inference attacks to audit unauthorized data use in RLVR training pipelines used in LLM post-training.
Large Language Action model (Humanoid-LLA) that translates free-form natural language commands to robotic humanoid control with diverse motions.
GraphBench provides standardized benchmarking suite for graph machine learning with consistent evaluation protocols for foundation models.
CARL improves multi-step reinforcement learning for agents by identifying and optimizing criticality-aware action choices rather than treating all steps equally.
Security analysis showing LLM-based tabular data generation systems memorize and leak sensitive string information from training data.
RAG-HAR applies retrieval-augmented generation with LLMs for human activity recognition without dataset-specific training.
Analysis of whether LLMs can estimate question difficulty and perceive student learning struggles for educational assessment.
Parameter-efficient fine-tuning method applying LoRA selectively to visual tokens and attention heads in vision-language models.
Framework (ThinkARM) for analyzing and abstracting LLM reasoning traces using episode theory into functional reasoning steps.
Framework aligning multimodal health sensor data with LLMs to generate clinical narratives for mental health assessment.