ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration
ChipCraftBrain: Multi-agent orchestration framework for LLM-based RTL code generation with validation-first approach and synthesis awareness.
ChipCraftBrain: Multi-agent orchestration framework for LLM-based RTL code generation with validation-first approach and synthesis awareness.
DR-Venus: 4B edge-scale deep research agent trained on 10K open data for cost-effective autonomous research task execution.
Mechanistic analysis of LLM quantization failure modes, revealing distinct mechanisms between 4-bit and 2-bit precision degradation.
LLM-based system for depression risk assessment in Reddit posts using multi-label classification on linguistic signals.
MMCORE: Unified multimodal framework using vision-language models to generate semantic embeddings for image diffusion models.
Data-driven ML framework for emission prediction and control in cement manufacturing using operational data from 4 plants.
Study examining whether LLM-powered agents systematically reflect behavioral characteristics of their human operators.
Empirical study reassessing correlation between membership inference attack success rates and model generalization.
4B-parameter vision-language model for wound infection classification with evidence-grounded clinical reasoning explanations.
DistortBench: Diagnostic benchmark evaluating vision-language models on image distortion identification across 27 types.
Semantic prompting: Interactive approach enabling agentic LLM-based narrative refinement through spatial semantic interactions.
Analysis of name-based bias in LLM-generated resume summaries across 4 models using synthetic and real-world data.
EmbodiedMidtrain bridges gap between vision-language models and embodied vision-language-action models through mid-training adaptation.
BMBE: Modular medical dialogue framework separating language models from probabilistic reasoning for autonomous diagnostic agents.
Fine-tuning Vision Transformers' self-attention on human saliency maps to align model attention with human visual perception patterns.
TriEx framework explaining multi-agent LLM reasoning through structured first-person self-reasoning and belief state visualization.
Empirical study of LLM agents aggregating private information through trading in prediction markets.
System for auditing and controlling AI agent actions in spreadsheets during execution with user oversight.
PLMA framework combining learning with warm-started MCMC for solving quadratic assignment problems.
Theoretical analysis of bilevel minimax optimization algorithms' generalization properties for hyperparameter optimization and reinforcement learning.
Adaptive conformal anomaly detection method using pre-trained time series foundation models for signal monitoring with interpretable p-value anomaly scores.
AgentSOC multi-layer agentic framework for SOC automation integrating perception, anticipatory reasoning, and risk-based response action selection for security alerts.
IMPACT-CYCLE multi-agent contract-based system for claim-level supervisory correction of long-video semantic memory with interpretable intermediate outputs.
Meta-Tool empirical study comparing hypernetwork-based LoRA adaptation vs few-shot prompting for tool-use in 3B small language models.
LLM-agent-based vulnerability detection system for Node.js packages using multi-step reasoning to overcome dynamic JavaScript analysis limitations.
Text-guided dual-gaze prediction model for interpretable driver attention using Vision-Language Models with scene-to-object semantic reasoning.
Characterization and benchmarking of logging code security vulnerabilities using LLMs to detect insecure logging practices and sensitive information exposure.
Hybrid Policy Distillation framework unifying knowledge distillation methods for LLM compression, reformulating KD as reweighted token-level log-likelihood objective.
Cortex 2.0 robotic system shifting from reactive Vision-Language-Action models to proactive world models for long-horizon manipulation tasks across embodiments.
Text steganography method using LLMs with dynamic codebook and multimodal models for hiding information in generated text.
Mobile GUI agent framework with adaptive visual modality selection for transparent human-agent interaction during smartphone task automation.
LLM used as experimental planner in closed-loop system for phase diagram construction of multicomponent alloys via high-throughput synthesis and X-ray diffraction.
Technical formalization of logit shifts induced by LoRA fine-tuning, decomposing multi-layer effects into layerwise contributions using Fréchet approximation.
Demonstrates that image generators develop strong zero-shot visual understanding capabilities through generative pretraining, similar to emergent LLM abilities.
Surrogate modeling framework to interpret and explain knowledge encoded in black-box LLMs for medical prediction tasks.
Vision-Language-Action model for automated ultrasound-guided needle insertion, replacing hand-crafted pipelines with learned adaptive control.
Using in-context learning with LLMs to enable bimanual robot manipulation without task-specific training, addressing coordination constraints in high-dimensional action spaces.
LaplacianFormer proposes Transformer variant using Laplacian kernel to reduce attention complexity for high-resolution vision tasks.
Research identifies hallucination in AI fluid dynamics models, where viscous fingering patterns appear realistic but are physically implausible.
CyberCertBench introduces benchmarks evaluating LLM domain knowledge against cybersecurity industry certification standards.
Research paper on Onyx, a cost-efficient disk-based approximate nearest neighbor search for sensitive data using oblivious RAM techniques.
Paper introduces Semantic Recall metric for evaluating approximate nearest neighbor search by filtering semantically irrelevant objects.
Performance analysis and optimization of BentoML-based AI inference system for scalable model serving and deployment.
Shift-Up: Framework establishing software engineering guardrails for AI-native development, addressing architectural drift and maintainability.
DialToM: Benchmark evaluating LLM Theory of Mind abilities through dialogue state prediction and functional utility assessment.
VTouch++: Bimanual robot manipulation dataset with vision-based tactile sensing for contact-rich tasks.
MOMO: Framework for robot skill learning through kinesthetic, natural language, and graphical interaction modalities.
Knowledge Capsules: Structured nonparametric memory units for LLMs enabling efficient knowledge updates without retraining or RAG limitations.
CHASM: First dataset for evaluating multimodal LLMs in detecting covert advertisements on Chinese social media.
Evaluates LLMs on feature model analysis for Software Product Line validation using semi-formal textual blueprints.