Multimodal Multi-Agent Empowered Legal Judgment Prediction
JurisMMA framework uses multimodal multi-agent systems to predict legal case outcomes by decomposing trial processes and handling multiple allegations.
JurisMMA framework uses multimodal multi-agent systems to predict legal case outcomes by decomposing trial processes and handling multiple allegations.
SA-SFT method mitigates catastrophic forgetting in LLM fine-tuning by generating self-dialogues and mixing with task data before training.
ConceptRM addresses alert fatigue in intelligent agents by using consensus-based data cleaning and reflection models to filter false alerts.
CAGE: Framework systematically adapting red-teaming prompts to cultural contexts for culturally-aware LLM safety evaluation and vulnerability detection.
Domain-specific LLM fine-tuned on physics-based energy simulations providing retrofit recommendations for residential buildings to bridge expertise gap.
MoBiQuant: Mixture-of-bits quantization enabling token-adaptive elastic LLM deployment across varying computational resources and precisions.
Demonstrates encoder-side poisoning induces trigger-free semantic corruption in diffusion models through Jacobian-based geometric analysis.
OpenPort Protocol: Governance-first specification for secure, least-privilege tool access in AI agents with auditability and abuse resistance.
Controllable exploration strategy for reinforcement learning with verifiable rewards in multi-modal LLMs to prevent entropy collapse and policy degradation.
Dual-memory augmented vision-language-action model improving inference efficiency and action generation for robotic manipulation tasks.
Framework automating forensic artifact extraction and validation using LLM analysis and Digital Forensic Knowledge Graph on forensic images.
Benchmark methodology for MLIR-based AI kernel compiler analyzing vectorization, multi-threading, and double buffering for edge device optimization.
Studies pedagogical guardrails needed when using LLMs for novice programming to avoid cognitive outsourcing and ensure skill acquisition.
Identifies optimal layers for knowledge editing in LLMs via layer gradient analysis to improve edit success rates without degrading performance.
ESM: Framework using principal component analysis on layer gradients to merge multiple task-specific models while reducing task interference.
CodeHacker: Automated agent framework generating adversarial test cases to expose vulnerabilities in LLM-generated code solutions.
Sovereignty kernel framework enabling tamper-evident, verifiable logging of AI agent execution for regulatory compliance and auditability.
KnapSpec: Training-free framework reformulating draft model selection as knapsack problem to optimize LLM inference throughput via adaptive layer skipping.
Multimodal framework combining vision-language models and speech processing with fuzzy logic for robotic arm control in human-robot interaction.
Systematic ablation study across 100 real-world robot runs identifying which design choices enable successful online reinforcement learning on physical robots.
Extends TabPFN foundation model to handle multimodal data (images, text, tabular) for healthcare and marketing applications.
Applies ConvexTopics clustering and LLMs to organize biomedical anti-aging literature and detect research trends.
Empirical study quantifying gaps between expected and actual performance of AI agents across software engineering, clinical documentation, and decision support domains.
Framework for evaluating LLM personality simulation using interview data from 671k+ interactions, grounding generation in authentic personal data rather than proxies.
Analysis of linguistic feature dimensions affecting LLM performance using 22-feature vector covering clause complexity, anaphora, and answerability.
PhysMem: memory framework enabling VLM robot planners to learn physical principles through interaction and test-time experience.
Circuit tracing framework for analyzing internal mechanisms of vision-language models using transcoders and attribution graphs.
QueryBandits: model-agnostic contextual bandit framework for hallucination mitigation in closed-source LLMs.
LLM-as-a-judge evaluation framework for enterprise RAG systems handling multi-turn case-based workflows with enterprise-specific failure modes.
Analysis of challenges in unsupervised elicitation and easy-to-hard generalization for steering language models toward truthful outputs.
Investigation of mechanisms reducing LLM idea diversity compared to humans, with analysis of fixation effects and potential mitigation strategies.
Analysis of how protein language models differ from natural language transformers and methods to improve inference on biological sequences.
Hybrid dialogue agent design combining rule-based pedagogy with LLM-based responses for learner reflection in educational settings.
Federated learning framework for multi-task LLM fine-tuning via sparse-orthogonal LoRA over wireless networks with heterogeneous datasets.
LESA: learnable predictors for accelerating diffusion transformer inference through feature caching adapted to diffusion process dynamics.
Apprenticeship learning framework for intelligent tutoring systems using reinforcement learning to overcome sample inefficiency and reward design challenges.
Actor-Curator: Automated curriculum learning framework using policy-improvement bandits for efficient RL post-training of large language models.
Study applying Technology Acceptance Model to understand what drives student adoption of conversational AI chatbots for learning.
R&R detector suite measuring memorization of personal information (emails, phone numbers, IPs) in language models trained on web data.
OptiLeak: RL-enhanced attack method exposing prompt leakage vulnerabilities in multi-tenant LLM services via shared Key-Value cache side-channels.
Comparative study of ML models (CNN, LSTM, BERT) for detecting hate speech and offensive language on social media using text transformation techniques.
TrajGPT-R: transformer model generating urban mobility trajectories using reinforcement learning and pre-training for privacy-preserving data synthesis.
CAMEL: confidence-gated reward model combining interpretability with efficiency, using log-probability margins for prediction confidence estimation.
PRECTR-V2: unified framework combining search relevance matching and CTR prediction with cross-user preference mining and LLM-distilled encoders.
UrbanFM: foundation model for urban spatio-temporal data analysis addressing generalization across regions and tasks in urban computing.
Agile V: compliance framework embedding AI agents for requirements, design, build, test, and deployment with verification and audit trails.
AdapTools: attack framework demonstrating indirect prompt injection vulnerabilities in agentic LLMs using external data services and tool integration.
Communication-inspired discrete image tokenizer for vision systems, optimizing for semantic structure and object-level representations rather than texture compression.
RMIT-ADM+S: RAG system with dynamic retrieval strategy routing based on query complexity, enabling efficient operation on consumer hardware.
SibylSense: inference-time learning approach adapting reward rubrics for open-ended generation tasks, addressing reward hacking and scaling rubric construction.