PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts
Parameter-efficient multi-task learning approach using optimized continuous prompts for fine-tuning a single LLM across multiple tasks.
Parameter-efficient multi-task learning approach using optimized continuous prompts for fine-tuning a single LLM across multiple tasks.
Speech-based benchmark dataset for early-stage Parkinson's disease detection with speaker-independent evaluation protocol.
Attention-guided training framework for interpretable genomic sequence classification using neural networks with saliency learning.
Technique for injecting constrained reasoning into code agents via nullspace editing, aligning planning capabilities with tool-use protocol discipline.
Edge-cloud architecture for automated diabetic retinopathy screening in rural healthcare settings using deep learning.
Agentic framework combining interpretable prototype networks with privacy-aware LLM workflows for clinical diagnosis documentation and explainability.
Text-based generative approach using LLMs and reinforcement learning with verifiable rewards to generate professional floor plans respecting numerical constraints.
Reinforcement learning approach for training tool-calling LLM agents to perform multi-step reasoning over FHIR healthcare data graphs.
Benchmark dataset for evaluating multilingual LLM safety in national security and public safety contexts across diverse language-geopolitical pairs.
Benchmark for evaluating LLM-based cybersecurity agents on exploitation tasks with granular capability levels, moving beyond binary success/failure metrics.
Method for improving RAG systems in dialogue assistants by using prospection-guided retrieval to recover semantically distant but contextually relevant facts from interaction histories.
Research study examining failure modes of RAG systems through a graph-based lens, analyzing how retrieved evidence influences LLM answer generation.
Empirical study of LLM-based robustness testing for microservice APIs, comparing model and prompt strategies for failure detection.
Fine-grained concept bottleneck models with visual grounding for interpretable predictions with verifiable concept evidence.
Prefill-only finetuning method enabling efficient personalized LLM serving without throughput degradation from user-specific adapters.
Identifies and diagnoses training-inference mismatch in LLM RL systems where rollout and optimization stages produce inconsistent token probabilities.
Active learning approach to improve pairwise ranking prompting from LLMs by reframing as efficient reranking problem.
Selective alignment knowledge distillation for spiking neural networks that weights timesteps unequally to improve performance over ANNs.
Analyzes spectral geometry of transformer residual streams across layers, revealing full eigenvalue distributions and coupling to network topology.
Watermarking techniques for game-playing agents in perfect-information games to detect unauthorized use and cheating in gaming platforms.
MetaMoE: Privacy-preserving framework for unifying independently trained domain-specialized experts into centralized MoE using public proxy data.
Web agents should use plan-then-execute paradigm instead of ReAct, committing to task-specific programs before observing web content to avoid injection attacks.
MMGuard: Proactive protection mechanism for multimodal data preventing unauthorized fine-tuning of vision-language models before training.
Policy optimization for hybrid discrete-continuous action spaces using mixed gradients to improve credit assignment in robotics and control.
MSRL: Geometric reinforcement learning that reuses local transition geometry via matrix descriptors for compositional generalization across tasks.
ICED: Concept-level machine unlearning for vision-language models using interpretable concept decomposition to remove specific knowledge.
Continuous semantic alignment for GUI agents: Improves test-time scaling by replacing binary critic classification with fine-grained ranking ability.
Dynamic Latent Routing: Language model post-training method using General Dijkstra Search for temporal composition of sub-policies in MDPs.
RQ-MoE: Dynamic vector quantization method using mixture of experts for efficient input-dependent compression of high-dimensional embeddings.
Context window filtering for LLM-based developer tools using correctness-aware repository filtering to maximize effective context within practical constraints.
LoMETab: Rank-r generalization of multiplicative ensembles for tabular deep learning that improves on gradient boosting and attention-based architectures.
DiHAL: A diffusion-transformer hybrid that identifies optimal layers for integrating continuous diffusion into pretrained language models for improved denoising.
Coherent Coordinate Descent optimization method for zeroth-order scenarios where backpropagation unavailable, improving sample efficiency and variance.
Optimal Pattern Detection Tree for symbolic rule-based classification providing interpretable rules for pattern discovery tasks.
Multi-agent starting-state sampling strategy to accelerate exploration in imperfect-information competitive games like StarCraft and Dota.
Darwin Family framework for training-free evolutionary merging of LLMs via gradient-free weight-space recombination to improve reasoning performance.
MARS framework for memory-augmented LLM agents in recommendation systems using structured belief-state memory with lifecycle management.
SWE-Chain benchmark for evaluating coding agents on realistic package upgrade chains, testing continuous maintenance tasks beyond isolated issue resolution.
Analysis of stochasticity problems in LLM jailbreak evaluation, showing reported adversarial attack methods fail to replicate promised performance against independent models.
MemLineage defense mechanism for LLM agent memory using cryptographic provenance and derivation lineage to prevent malicious state injection into agent sessions.
Federated actor-critic framework where agents share linear subspace representation while maintaining personalized local policies for collaborative training.
Generative retrieval framework for e-commerce search with semantic cluster IDs and expert-guided RL, optimized for industrial deployment.
Single-pass framework (QAOD) for hallucination detection in LLMs using question-answer orthogonal decomposition without repeated inference.
Framework for detecting when RAG systems inappropriately rely on retrieved context conflicting with model knowledge, with inference-time intervention.
Diagnostic study on retrieval-augmented code generation showing how temporally stale repository context degrades completion quality.
Multi-agent framework with arena-based argumentation for transparent multimedia verification using multimodal LLMs and verification tools.
Offline-to-online reinforcement learning approach using bi-level optimization for adaptive data mixing across training stages.
Deep reinforcement learning method for dynamic rebalancing of dockless bike-sharing systems with real-time truck routing.
Framework for dimension-level intent fidelity evaluation of LLMs through structured prompt ablation across multiple languages and domains.
Prescription-level benchmark (RxEval) for evaluating LLM medication recommendation systems with temporally-aware medical decision-making.