Multi-objective Evolutionary Merging Enables Efficient Reasoning Models
Multi-objective evolutionary merging approach to reduce computational overhead of reasoning models while maintaining accuracy with fewer tokens.
Multi-objective evolutionary merging approach to reduce computational overhead of reasoning models while maintaining accuracy with fewer tokens.
Practical implementation of activation-level interpretability and steering techniques for large language models distributed across multiple GPUs.
Symbolic Equivalence Partitioning uses symbolic execution for inference-time code selection in LLM-based code generation without expensive verifiers.
DoMinO framework unifies reinforcement learning fine-tuning of Discrete Flow Matching models by reformulating sampling as a multi-step MDP.
MedConclusion benchmark dataset of 5.7M PubMed abstracts for evaluating LLMs on biomedical conclusion generation from structured evidence.
Efficient quantization method for Mixture-of-Experts models with theoretical generalization guarantees to reduce inference memory overhead.
Soft-Quantum Algorithms explores classical simulation of variational quantum circuits for few-qubit problems with large datasets.
SkillSieve is a three-layer detection framework for identifying security vulnerabilities in AI agent skills, addressing both code and natural language prompt injection attacks.
AI-Driven Research for Systems uses LLMs to automate database performance optimization through automated code generation instead of manual design.
Guardian Parser Pack uses LLMs to parse and normalize heterogeneous investigative documents for missing-person cases with varying layouts and data quality.
SciDC method reduces LLM hallucination by incorporating scientific knowledge and rules as decoding constraints to improve reliability.
TwinLoop framework uses simulation-in-the-loop digital twins for online multi-agent reinforcement learning to adapt policies when operating conditions change.
Research finding that 52-88% of chain-of-thought tokens in reasoning models are generated after the answer is already recoverable, revealing a detection-extraction gap in model behavior.
CubeGraph: efficient retrieval-augmented generation system for hybrid queries combining vector similarity search with spatio-temporal filters for RAG workloads.
Logical Robots: declarative multi-agent programming platform using logic programming language Logica for robot behavior specification combining reactive control and planning.
SubFLOT: Federated learning method using optimal transport for efficient submodel extraction, addressing heterogeneity and enabling client-side personalization.
SHAPE: Framework for improving LLM reasoning through process supervision, formalizing reasoning as state-space trajectory with stage-aware advantage estimation.
Parameter-efficient multitask prompt distillation framework for clinical NLP adapting shared metaprompts across diverse medical tasks.
Audience segmentation approach for LLM-based social simulation restoring demographic heterogeneity in behavioral modeling.
Fake news detection framework combining graph analysis with LLM-retrieved evidence for explainable veracity assessment.
Chemical vision-language model emphasizing reasoning over perception for understanding molecular reactions and mechanisms.
Confidence calibration methods for LLM-generated code revisions enabling developers to assess output correctness at instance-level.
Open-source Chinese legal language model built on Baichuan foundation using continued pretraining and instruction tuning.
CLI-Tool-Bench benchmark for evaluating LLM agents' end-to-end software generation from intent without predefined scaffolds.
Framework enabling multi-LLM collaboration with role-based team structure for solving complex multi-step contextualized tasks.
Pipeline for extracting procedural knowledge and directed graphs from maintenance flowchart images using vision-language models.
Multi-faceted preference alignment approach for conversational query rewriting using feedback from retrieval and generation components.
Benchmark for evaluating LLM-generated repository documentation using question answering, addressing limitations of LLM-as-judge evaluation methods.
FedDAP addresses domain shift in federated learning using prototype learning for privacy-sensitive applications.
Instance-adaptive variational autoencoders reduce amortization gap in latent variable models for deep generative modeling.
MoBiE: binarization framework for efficient inference of mixture-of-experts LLMs using post-training quantization.
SkillTrojan: backdoor attack framework targeting skill-based agent systems through malicious skill implementations.
OmniTabBench: largest tabular data benchmark comparing GBDTs, neural networks, and foundation models at scale.
ESG sentiment analysis dataset and models for Slovene news, addressing corporate performance assessment in emerging markets.
WRAP++ improves LLM pretraining through synthetic data rephrasing that captures cross-document relationships and associative context.
Privacy-preserving LLM inference method enabling text-free processing through alignment and adaptation, reducing privacy risks without computational overhead.
Analysis of step length confounding bias in LLM reasoning dataset selection pipelines used for fine-tuning complex reasoning models on chain-of-thought tasks.
Memory architecture for long-term dialogue systems using boundary-guided event segmentation and query-adaptive retrieval to improve scalability and personalization.
Benchmark for evaluating LLM diagnostic robustness in medical dialogue with adversarial patient behaviors at varying severity levels and cross-dimension interactions.
Large-scale study examining bias in skin-toned emoji representations across LLMs and embedding models, addressing societal bias perpetuation in AI systems.
Research on redundancy in Large Speech Language Models reveals structured token-level redundancy to reduce inference costs while maintaining semantic fidelity.
SentinelSphere combines ML-based threat detection with LLM-powered security training to address cybersecurity skill gaps and human vulnerabilities.
Extended reality platform combining XR and multimodal AI for personalized career guidance and coaching.
Benchmark measuring occupational skill susceptibility to LLM automation across 263 tasks and 35 O*NET skill categories.
Query-aware adaptive perception method for reducing tokens in multimodal LLM inference while maintaining fine-grained understanding.
Efficient scaling method for diffusion model reinforcement learning using mixed precision and selective rollout quantization.
Multi-modal UI control detection combining YOLO vision with GPT-generated text descriptions via cross-attention.
Neural improvement method learning local search policies for TSP, generalizing beyond single-solution outputs.
Empirical study of LoRA fine-tuning LLMs for automated test case generation from natural language requirements.
Study of self-preference bias in LLM-as-judge evaluation, showing models favor outputs from themselves or related models.