Open-source privacy firewall for LLM interactions combining browser extension and proxy. Intercepts HTTP/S and WebSocket traffic for data protection.
Multi-agent pipeline using MCP for converting legacy documentation to NIST OSCAL compliance format. Non-invasive threat assessment for critical infrastructure.
Git-based versioning system for data lakehouse enabling agents to work on isolated branches with human review. Production system with atomic publishing.
Analyzes type information encoding in pre-trained code models through hidden state probing. Shows cross-lingual type representations emerge from untyped code.
Fast-Slow dual-system for UAV vision-language navigation balancing semantic reasoning with real-time flight control.
Studies timing properties in synthetic conversational speech data as controllable variable for training conversational ASR systems.
Self-adaptive anomaly detection system using reinforcement learning and human feedback for connected vehicle monitoring.
Uses LLMs as judges to align metric learning in personality recognition, moving beyond theory-dependent taxonomies.
Dual-level Vision-Language-Action model combining semantic forecasting and world evolution for proactive autonomous driving.
Watermarking scheme for LLM-agent trajectory logs enabling attribution verification against resellers with full data access.
Disease-aware generative language model for drug discovery that conditions molecular generation on disease ontology and protein sequences.
Uses Group Relative Policy Optimization to adapt LLM-based ASR systems trained on synthetic speech for real-world banking domain.
World Action Models trained on egocentric human video for robot manipulation, separating transferable task semantics from human-specific factors.
Q-learning approach for adaptive model retraining in Open RAN networks to handle traffic-induced performance drift.
Research on LLM abstention mechanisms distinguishing between incorrect answers and unanswerable questions using separate confidence axes.
Research paper analyzing interaction-level disparities in AI agent access beyond availability/quality/quantity dimensions.
Multimodal agent with episodic memory for understanding, generation, and editing without context window explosion in long-horizon dialogue.
Audits reliability variation in LLM-as-judge evaluation across model upgrades, showing judge replacements are not interchangeable measurement tools.
DocMaster hierarchical document analysis system preserving structural relationships for LLM-based analysis of academic papers and technical documents.
VocaDet open-vocabulary object detection and segmentation via visual tokenization and vector database retrieval for scalable category expansion.
SMetric session-centric LLM scheduling system optimized for agentic serving with high KV-cache reuse, balancing throughput and latency.
Structured sparse autoencoders for learning modality-consistent concepts across vision-language models with improved mechanistic interpretability.
UltraX adaptive programmatic editing system for large-scale pre-training data refinement beyond rule-based and rule-learning approaches.
Multi-modal machine teaching approach for robust reward learning in autonomous agents across diverse operational environments using inverse RL.
WebSwarm multi-agent orchestration system for deep-and-wide web search using recursive agent coordination beyond single trajectory limitations.
Training-free speculative decoding acceleration for LLM sampling with relaxed distribution guarantees enabling speed-capability trade-offs.
ProjAgent retrieves procedurally similar repository functions for repository-level code generation, accounting for cross-file dependencies and project conventions.
Validity assessment of Portugal's AMALIA 9B LLM for data annotation tasks, comparing agreement and reliability against human coders.
In-training low-rank regularization technique for neural network compression without requiring SVD or architecture modification.
Chain-of-Frame reasoning approach using video generation models as alternative to chain-of-thought for logical reasoning in large models.
Multi-perspective causal discovery framework using LLMs for abductive reasoning, includes DeepAbduction dataset for pollution cause analysis.
Curriculum learning method for Direct Preference Optimization using two-dimensional difficulty space (prompt complexity and distinguishability) to improve LLM alignment.
Research on energy-efficient domain-specific AI models and agents, addressing computational costs of large language models in production.
System for training task-oriented dialogue agents for recruitment using simulator-based data evaluation and selection.
Case study using ChatGPT for rapid scientific prototyping in lunar trajectory estimation competition, achieving second place.
Benchmark for evaluating long-horizon autonomous decision-making and strategy stability of LLM agents in supermarket simulation.
Position paper arguing RAG systems need redesign to handle opinion-rich content beyond factual grounding.
Neural architecture for learning long-range non-stationary temporal patterns in streaming settings without revisiting past data.
Study of memory design in long-lived foundation model agents, analyzing personalization, extraction risk, and deletion fidelity tradeoffs.
Lightweight LLM-based agent framework for rare disease diagnosis built through policy iteration with human feedback.
Tool for formal verification of neural ordinary differential equations in safety-critical applications.
Multi-layer benchmark with 118 problems for evaluating LLMs on reconstructing and applying expert investment decision frameworks.
System prompt technique (narration-of-thought) for improving LLM ethical reasoning on moral dilemmas by reducing stakeholder collapse.
Token-efficient context management module for LLM agents performing repository-level program repair with precision evidence selection.
Analysis of security vulnerabilities in persistent-state AI coding agents that ship code iteratively across sessions.
Benchmark for evaluating LLM agents across multilingual long-horizon tasks requiring planning, tool use, and environment interaction.
Theoretical framework for embedding cognitive architectures natively into LLMs rather than simulating via prompting and context management.
Expert-based analysis of explainable AI methods for safe AI development and certification under regulatory frameworks.
ParamMute method suppresses knowledge-critical feed-forward networks to improve faithfulness in retrieval-augmented generation systems.
Synthetic data generation framework for personalizing vision-language models using concept hierarchies.