TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models
Framework enabling small reasoning models to accelerate large reasoning model inference by detecting capability boundaries and reasoning risks.
Framework enabling small reasoning models to accelerate large reasoning model inference by detecting capability boundaries and reasoning risks.
Trajectory-level safety benchmarks for AI agents in OpenClaw and Codex environments with diagnostic evaluation tools.
Framework addressing LLM reliability issues in decision-making by adding knowledge layer to detect reasoning drift and uncertainty.
Federated learning framework for cross-silo healthcare collaboration with incentive mechanisms for coopetitive data contribution.
MemoSight: unified framework combining context compression and multi-token prediction to accelerate chain-of-thought reasoning in LLMs.
Agentic RAG system for Ukrainian combining two-stage retrieval with lightweight agentic query rephrasing and retry loops on small LLM.
Framework for human-AI collaboration treating reasoning as distributed relational process with epistemic scaffolding and traceable reasoning.
ADAPT/DynAfford benchmark evaluating embodied agents' ability to plan under object affordance constraints and unspecified environmental conditions.
Dual-axis generative reward model for training full-duplex spoken dialogue agents with semantic and turn-taking robustness via reinforcement learning.
WavAlign: adaptive hybrid post-training method enhancing open-source spoken dialogue models via online reinforcement learning.
Task-capability coevolution framework discovering novel LLM experts with emergent diverse skills through open-ended model-task evolution.
Conformal learning approach for hybrid decision-making where vision-language models provide guidance with uncertainty calibration for human decisions.
Dr.RTL: autonomous agentic system for realistic RTL circuit optimization using tool-grounded self-improvement with fine-grained modifications.
COEVO: co-evolutionary framework optimizing functional correctness and PPA simultaneously in LLM-based RTL hardware code generation.
Mixture-of-experts flow matching framework for faster language model inference while maintaining generation quality.
Autogenesis Protocol: self-evolution protocol for LLM-based agent systems addressing lifecycle management, version tracking, and safe updates.
ProVoice-Bench: evaluation framework for proactive voice agents with four novel tasks assessing multimodal interaction beyond reactive responses.
Scoping review synthesizing research on fairness in multi-agent AI systems across 23 studies, identifying five archetypal fairness approaches.
OpenMobile: open-source framework for mobile agents using vision-language models, with task and trajectory synthesis for Android task automation.
Open-source HyperSpace framework decomposing Vector Symbolic Architecture systems into modular operators for hyperdimensional representations.
Vector Symbolic Architecture method for sequential associative memories in streaming environments with imbalanced, non-stationary data.
arXiv paper on SRMU, hyperdimensional memory architecture for streaming associative memories with relevance-gated updates.
arXiv paper on axiomatic benchmark for evaluating scientific novelty metrics with AI involvement in research generation.
arXiv paper on IG-Search reinforcement learning framework for training LLMs with step-level rewards in search-augmented reasoning.
arXiv paper on agentic systems for iterative CAD model design using code generation and visual feedback loops.
arXiv paper on policy-guided dual-process user simulation for merchant behavior analysis and counterfactual evaluation.
arXiv paper on IRS framework for humor understanding with structured reasoning supervision on multimodal datasets.
arXiv paper exposing vulnerability in LLM-as-judge evaluation where contextual framing influences assessment outcomes.
Blue Data Intelligence Layer: Agent-based system for NL2SQL queries spanning multiple data sources with streaming data and multimodal inputs.
Interpretability study examining how LLMs and VLMs understand viewpoint rotation using only text inputs without visual information.
Diagnostic toolkit for LLM-as-judge reliability using conformal prediction and transitivity analysis revealing widespread per-instance inconsistency.
Controlled study of LLM generalization using shortest-path planning to isolate effects of training data, paradigms, and inference strategies.
Edge-cloud collaborative architecture for elderly care using real-time risk assessment and emergency response systems.
MemGround: Benchmark for evaluating long-term memory capabilities of LLMs in gamified interactive scenarios with dynamic state tracking.
HUOZIIME: On-device LLM-powered input method editor enabling personalized text generation with privacy preservation on mobile devices.
Study investigating whether LLMs can identify methodological flaws like data leakage in published ML research papers through automated analysis.
SeaAlert: information extraction from maritime distress communications using LLMs for emergency response.
LoRA fine-tuning and in-context learning for Chinese rhetoric recognition in automated essay scoring.
SAGE Celer 2.6: general-purpose language models (5B-27B) with inverse reasoning pipeline for validation and hallucination reduction.
Stateful evidence-driven RAG framework with iterative reasoning for grounding LLMs in external knowledge.
Benchmarking linguistic adaptation of Llama-3.1-8B, Mistral-7B, and Qwen3-8B on Romanized Nepali language.
Teacher-guided retrieval-augmented generation framework for resolving conflicting vulnerability information in LLMs.
Two-stage QLoRA fine-tuning of Qwen3-4B for clinical question answering and evidence sentence alignment task.
Energy-aware gradient coordinator addresses gradient entanglement in generalized category discovery optimization.
SPFG: dataset and task for generating spoken pedagogical feedback with grammatical error correction and explanations.
Scoping review of LLM applications for rare disease patient education and communication support.
318M parameter Transformer trained on Classical Chinese: investigation of internal knowledge vs external expression in language models.
HARNESS: lightweight distilled Arabic speech foundation models trained with iterative self-distillation.
Evaluation of LLMs on tacit reasoning in quantum field theory and string theory with non-binary correctness metrics.
PICCO framework: taxonomy and reference architecture for structuring LLM prompts based on 11 published frameworks.