RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
Proposes RRCM framework for LLM-based recommender systems using ranking-driven retrieval over collaborative and meta memories.
Proposes RRCM framework for LLM-based recommender systems using ranking-driven retrieval over collaborative and meta memories.
Constructs Adversarial Empathy Benchmark to probe robustness of RL-trained empathetic language agents against adversarial user interactions.
Proposes Distillation through Reasoning Path Compression to improve consistency of teacher rationales when distilling reasoning into smaller LLMs.
Introduces MathlibPR benchmark for evaluating LLM-assisted pull request merge-readiness in Lean formal mathematics library.
Audits text embeddings using citation graphs of 3.58M papers, revealing disconnect between cosine similarity and conceptual relatedness in RAG.
Benchmarks attention transfer across 20 Vision Transformer teachers, finding it not universally effective across ViT families.
Develops closed-form dataset distillation method for linear probing on frozen pre-trained vision model encoders.
Proposes Three-in-One world model combining Deep Boltzmann Machine with task-specific heads for marketing intervention modeling.
System for toxicity detection in gaming chat using fine-tuned LLMs with LoRA and synthetic data augmentation across six toxicity classes.
Introduces proxy-analyzer framework using open-weight models to detect hallucinations in LLM outputs via internal activations.
Proposes Planning-after-Trial adaptive policy for test-time code generation that efficiently allocates compute based on problem difficulty.
Presents MIPIAD defense framework against multilingual indirect prompt injection attacks in RAG and tool-using LLM systems using ensemble learning.
Introduces structured role-aware policy optimization for multimodal reasoning in large vision-language models using reinforcement learning from verifiable rewards.
Proposes sparse random-feature neural networks with Krylov-based SVD for solving singularly perturbed ODEs with improved scalability.
Analyzes generalization bounds for trained Transformers using spectrum-adaptive methods to improve upon existing norm-based complexity bounds.
DoLQ: method for discovering ordinary differential equations combining symbolic regression with LLM-based qualitative evaluation.
Sparse autoencoder architectures (Crosscoders, Diff-SAE) for detecting backdoor attacks in language models via mechanistic interpretability.
Analysis showing mean-pooled cosine similarity is length-dependent; proposes length-invariant alternative for comparing neural representations.
Memory-efficient alternatives to error correction codes for protecting deep learning models against hardware faults.
Sparse autoencoders as lightweight detection mechanism for adversarial attacks on vision-language models.
Data selection framework for large multimodal models using incremental optimization utility ranking instead of LLM-as-Judge.
Reinforcement learning approach for distilling compact GUI agents that work on-device across platforms.
Conformal prediction framework for uncertainty quantification in object detection with finite-sample coverage guarantees.
Theoretical analysis of sample complexity in contrastive representation learning with dependent tuple sampling.
Framework for evaluating safety vs. capability in phone-use agents, addressing ambiguity in existing benchmarks.
Study showing post-training reduces LLM alignment with human behavior; introduces Psych-201 dataset for measuring behavioral alignment at scale.
Decentralized multi-agent pathfinding solver using learned local communication for scalable trajectory planning in robotics and logistics.
MAVEN multi-agent verification network with epistemic auditing for reasoning tasks, enabling intermediate verification in CoT traces.
Prefix consistency method for improving Chain-of-Thought reliability by using answer reproduction as verification signal.
Data valuation method using quotient semivalues to resist false-name manipulation in ML data attribution.
Game-theoretic analysis of differential privacy auditing when audited developers can strategically respond to audit queries.
FactoryBench benchmark evaluates time-series models and LLMs on industrial robotic telemetry along four causal levels.
Memory-efficient recurrent LLM architecture decoupling compute from memory in looped reasoning models with sublinear KV cache.
Research on LLM self-assessment using cognitive appraisal theory as alternative to confidence scores for performance prediction.
CADTestBench introduces first test-based evaluation benchmark for Text-to-CAD task using automated testing methodology.
arXiv: MatryoshkaLoRA - adaptive rank LoRA for efficient LLM fine-tuning without grid search. Parameter-efficient training.
arXiv: Quantum-inspired optimization for non-convex ML problems in high dimensions. Theoretical optimization approach.
arXiv research: Tool-calling in LLMs is linearly readable/steerable via internal activations. Tested 12 models, 77-100% steering accuracy.
First-order optimization methods for bilevel problems with minimax structures in upper and lower levels.
Global training method for Spiking Neural Networks via parameter reconstruction addressing surrogate gradient approximation errors.
STARFlow2 unifying language models and normalizing flows for multimodal text-image generation with autoregressive normalizing flows.
PET-Adapter framework for test-time domain adaptation in medical image reconstruction handling Poisson noise and limited-angle acquisitions.
DR-ME test for interpretable distributional treatment effects, first semiparametrically efficient finite-location test for detecting distributional shifts.
Byte Latent Transformer (BLT) addressing slow byte-level autoregressive generation with diffusion-based training and generation techniques.
Mathematical study of non-negative L1-approximating polynomials under Gaussian distributions with applications to computational learning theory.
Normalizing Trajectory Models (NTM) for efficient diffusion-based generation with few steps while preserving likelihood framework.
Multi-stage prototype learning framework for interpretable multivariate time series classification identifying predictive temporal patterns.
Theoretical analysis connecting contrastive learning data augmentation to positive-incentive noise estimation via information theory.
UNA framework unifying diverse feedback types (preferences, scores, scalars) for efficient LLM alignment across RLHF and DPO methods.
Algorithm for testing whether training data satisfies noise model assumptions in computational learning theory.