Mesh Based Simulations with Spatial and Temporal awareness
Advances machine learning surrogates for CFD using GNNs and transformers, addressing training paradigm bottlenecks in physics simulations.
Advances machine learning surrogates for CFD using GNNs and transformers, addressing training paradigm bottlenecks in physics simulations.
Proposes agentic AI-native 6G networks where LLM-based agents operate as bounded, policy-governed entities for autonomous intelligence.
Autonomous multiagent framework automating mechanistic interpretability of LLMs through iterative hypothesis testing and feature discovery agents.
Neuro-symbolic agents framework combining LLMs with formal methods to enable flexible requirements reuse while maintaining structural validity and consistency.
Thesis on model merging paradigm: combining independently trained neural networks in weight space without optimization or original training data access.
SkillGraph-Service: hybrid microservice merging knowledge graphs with LLM fallback for competency framework interoperability and skill search.
Introduces S²R², a segment-level robustness framework for LoRA-tuned language models that enforces consistency at entity and relation levels beyond whole-sequence.
Tests causal inner products for cross-lingual concept transport in transformer representations across 17 models, finding anti-concentration phenomena.
Examines how agentic AI interfaces should prioritize explanation and oversight communication over routine user interaction during workflow execution.
Introduces Prosa, a Brazilian Portuguese LLM evaluation benchmark using rubric-based scoring with multi-judge filtering to reduce judge model bias.
Studies AI alignment through law-and-economics models of deterrence, analyzing how agentic AI systems respond strategically to incentive structures.
Brain-inspired spiking neural network with oscillatory dynamics and time-delayed coordination for cognitive learning mechanisms.
TRIMMER uses self-supervised reinforcement learning for domain-agnostic video summarization without manual annotations.
IMPACT-HOI mixed-initiative framework for egocentric video annotation of Human-Object Interactions to generate robot manipulation training data.
IMPACT-Scribe framework for dense video annotation using interactive correction-driven collaboration between human annotators and models.
MissBGM Bayesian generative model for missing data imputation with uncertainty quantification using neural networks and Bayesian inference.
Adaptive differential privacy method for fall detection in sensor data that applies class-aware noise to preserve prediction performance.
GRAVITY architecture for long-horizon conversational agents that structures retrieved memory fragments with relational and temporal anchors.
LLM-based adaptive agent for extracting information from BIM models through iterative exploration instead of static query translation.
Technique to surgically remove memorization signatures in unlearned LLMs using cross-sequence probes without capability loss.
SplitZip compression method for KV cache transfer in disaggregated LLM serving, optimizing prefill-decode architecture bottleneck.
TCDA model for conversational sentiment analysis using graph networks and positional encoding to handle multi-turn dialogues with temporal context.
Training-free method to mitigate object hallucination in vision-language models by steering caption generation.
Security analysis of agentic-AI runtimes showing critical gaps in action auditing and safety properties across tool calls and device actuation.
Research on dynamic grounding failures and repair mechanisms in multi-agent LLM negotiation across conversational turns.
Study of the Compliance Gap where AI agents verbally agree to instructions but fail to follow them, examining the disconnect between stated and actual behavior.
LLM-guided evolutionary search (AlphaEvolve) to discover closed-form solutions for radar power allocation in multi-target tracking.
Selector-guided curriculum learning approach for one-shot reinforcement learning from verifiable rewards in LLM math reasoning.
Benchmark (RMGAP) for evaluating reward model generalization across diverse user preferences in RLHF-aligned language models.
Framework for remote control with communication-constrained channels where actors receive action guidance from a centralized controller.
Evaluation of data-poisoning backdoor attacks on contrastive learning models and their generalization across datasets.
Analysis of hidden-state dynamics in large reasoning models to determine if extended reasoning traces reflect genuine internal computation or verbosity.
Research on exploration budget allocation in cooperative multi-agent reinforcement learning using intrinsic motivation and novelty bonuses.
Method for selecting optimal training subsets under label noise by leveraging data symmetries, detecting low-quality examples via k-NN based analysis.
Chart-FR1: Multimodal LLM benchmark for fine-grained reasoning on high-information-density charts with multiple subplots and dense annotations.
SANTA: Sparse attention method for memory-bound LLM inference that samples from post-softmax distribution to reduce KV cache bandwidth at long contexts.
RefusalGuard: Geometry-preserving fine-tuning method maintaining safety alignment in LLMs by analyzing activation-space feature changes during task-specific adaptation.
Unified benchmark (PepSpecBench) for peptide tandem mass spectrometry prediction addressing evaluation challenges in deep learning for computational proteomics.
Phone2Act: Hardware-agnostic teleoperation framework using smartphone and ARCore for low-cost, scalable vision-language-action data collection across robot platforms.
TRAP attack targeting world-model planning agents through tail-aware ranking manipulation to compromise long-horizon planning and generalist agent behavior.
Trojan Hippo attack exploiting LLM agent memory systems for data exfiltration through dormant payloads planted via single untrusted tool interactions.
First large-scale reproducible benchmark for machine learning on Raman spectroscopy, standardizing evaluation across fragmented datasets and spectral models.
AI agent system for automating legal judgment drafting using agentic information retrieval and rubric-guided generation to improve evidence recall and reasoning.
Training-free LLM approach with prompt engineering for conventional commit classification, eliminating need for large labeled datasets and ML models.
Multimodal dataset addressing ambiguity resolution in machine translation by combining visual grounding with text, identifying data quality issues in prior benchmarks.
Low-cost modular robotic platform (VILAS) for vision-language-action policy learning and deployment with integrated arm, gripper, and dual-camera perception.
Multi-variant reliability audit of 15 open-weight language models showing single-prompt accuracy misses important failure modes in calibration and robustness.
Framework for standardizing AI evaluation through randomized controlled trials, adopting validity principles from established experimental disciplines.
Coopetition-Gym v1 benchmark platform for mixed-motive multi-agent RL with twenty environments covering interdependence, trust, collective action, and sequential dynamics.
Pair2Score transfers pairwise LLM comparisons into absolute essay scoring predictions using parameter-efficient LLaMA adaptation in two-stage framework.