Policy-Governed LLM Routing with Intent Matching for Instrument Laboratories
Policy-governed LLM routing system (Routiium) for engineering lab assistance with instructor control over timing and cost.
Policy-governed LLM routing system (Routiium) for engineering lab assistance with instructor control over timing and cost.
Empirical study on moral perception of AI-generated content reuse examining authorship and plagiarism concerns.
Studies validity of multimodal LLM feedback on student hand-drawn scientific diagrams and visual model encoding.
Proposes ethical emotion regulation framework for agentic AI inspired by 17th-century philosophy for autonomous learning environments.
Multi-agent guardrails system for clinical safety and hallucination mitigation in patient-facing LLM healthcare applications.
Study of systematic biases in transformer-based agentic AI systems deployed for user-facing applications like shopping and content navigation.
Language models with dataflow-aware pretraining for static program slicing, addressing hallucination and dependency modeling in code analysis.
DeepTutor: Open-source agentic framework for personalized tutoring using LLMs with adaptive feedback beyond RAG systems.
Educational framework using explainable AI and 20 Questions game format for adaptive cybersecurity training with interactive learning.
Analyzes prevalence and impact of AI-generated and AI-edited text on the internet using representative sampling methodology.
Addresses KV cache memory bottleneck in large-scale GPU inference serving through predictive multi-tier memory management and unified cache sizing.
Multi-agent system with self-evolving skill hub for optimizing large-scale recommendation system pipelines across pre-ranking, ranking, and re-ranking stages.
Proposes hierarchical framework for knowledge graphs that autodiscovers adaptive decay rates for different knowledge types instead of uniform forgetting curves.
Proposes self-conditioning adaptation for masked diffusion models to improve cross-step refinement in discrete sequence generation through iterative denoising.
Agent Name Service: DNS-inspired trust layer for Kubernetes providing secure discovery, cryptographic authentication, capability attestation, and policy governance for AI agents.
Analysis of continual learning challenges in memory-augmented LLM agents, showing stability-plasticity dilemma resurfaces at memory retrieval level under context limitations.
Empirical study of LLM performance variability and consistency in study screening for systematic literature reviews compared to classical methods.
FairMind: AutoML framework automating fairness analysis at dataset level using causal assumptions and LLM-generated reporting.
Neurogenesis-based approach for continual learning that dynamically expands network capacity without oracle task information.
Dual-stream memory architecture for persistent LLM health coaching agents reconciling patient self-report with electronic health records.
Study addressing learning rate transfer in normalized transformers (nGPT) across model dimensions using alignment exponents.
RoundPipe: pipeline parallelism technique with CPU offloading for efficient LLM fine-tuning on consumer-grade GPUs with limited memory.
CarryOnBench interactive benchmark measuring LLM ability to recover helpfulness through user intent clarification in multi-turn conversations while maintaining safety.
Multi-agent federated reinforcement learning system for autonomous vehicle lane change advisory with priority-aware decision making.
Knowledge distillation approach reducing SAM 3 and DINOv3 foundation models from 446M to 40.66M parameters for edge deployment in livestock monitoring.
Research on improving Linux privilege escalation capabilities of locally-hosted open-weight LLM agents for autonomous penetration testing.
Reformulation of guidance in generative modeling as optimal control problem with efficient algorithms for few-step reward-guided sampling.
Method for explaining uncertainty in conformal prediction by localizing calibration sources to distinguish aleatoric and epistemic uncertainty at instance level.
Mechanistic study of why LLM agents deviate from Nash equilibria in game theory tasks, using interpretability methods on Llama-3 and Qwen2.5 models to identify and reverse deviation causes.
Architectural approach separating explicit thinking and no-thinking modes in hybrid reasoning LLMs through Path-Lock Expert mechanism to reduce reasoning leakage.
Research on using LLMs for research software development where specifications evolve, identifying hallucination accumulation and unsupported assertion propagation as key failure modes.
Study on how freelance knowledge workers use generative AI tools for skill acquisition and the challenges they face in rapidly evolving skill demands.
Framework addressing adoption challenges for agentic AI systems in education, balancing implementation feasibility, personalization benefits, and disruption risks.
Research on adversarial instruction evaluation showing how LLMs fall back on positional shortcuts under high instruction complexity rather than engaging with question content.
Study on reasoning controllability in LLMs examining whether fundamental reasoning patterns like induction, deduction, and abduction can be decoupled from specific problem instances.
Paper introducing self-evolving software agents combining BDI reasoning with LLMs to enable autonomous evolution of goals, reasoning, and executable code.
Research on threat modeling for LLM-integrated robotic systems, analyzing how compromised inputs and unsafe outputs propagate through planning pipelines to physical consequences.
Studies serialization friction when LLMs process 2D structured tasks as 1D token sequences, impacting performance.
Examines epistemic guardrails in LLM reading assistants to prevent interpretive displacement and unsafe meaning-making.
Risk-sensitive contextual bandit approach for memory retrieval in LLM-based coding agents with safety constraints.
BoostLoRA gradient-boosting framework for parameter-efficient fine-tuning that grows adapter effective rank iteratively.
Pragmos system uses LLMs and agents for automated business process model derivation from textual descriptions.
End-to-end LLM framework for automated security operations combining threat detection, query generation, and resolution.
Survey of pre-service teachers' adoption intentions for AI-enabled educational tools using UTAUT2 framework.
Study of AI dependency among Filipino college students and effects on critical thinking and academic skills.
TypeBandit method for attribute completion in heterogeneous graph neural networks via type-level context allocation.
COHERENCE benchmark for evaluating multimodal LLMs on fine-grained image-text alignment in interleaved document contexts.
Applies Reliable Change Index from psychology to evaluate LLM version improvements on MMLU-Pro, finding most items show no reliable change.
Security research demonstrating secret extraction from local LLM fine-tuning through compromised model code backdoors.
AdaBFL framework for Byzantine-robust federated learning against poisoning attacks through adaptive aggregation.