GLM-5 foundation model transitioning from vibe coding to agentic engineering with DSA cost reduction and async RL infrastructure for improved autonomy.
Study of representation collapse during neural network training across five model scales showing scale-invariant emergence patterns in 119 task combinations.
AI-CARE metric for evaluating ML models on carbon emissions and energy consumption alongside standard performance metrics.
Research on interpretable Graph Neural Networks using symbolic methods to overcome message-passing limitations and Weisfeiler-Lehman expressivity barriers.
Neuro-symbolic framework (NSGGM) for molecule and graph generation combining neural proposals with symbolic guarantees for controllable generation.
MPZCH indexing mechanism for large-scale recommendation systems to mitigate embedding collisions and improve model freshness in embedding tables.
MASPO algorithm for LLM reasoning via reinforcement learning, addressing gradient utilization, probability mass, and signal reliability in trust region mechanisms.
Theoretical framework for training modular LLMs by combining domain-specific experts without heuristic dataset weighting, matching monolithic model performance.
Framework for multi-round human-AI collaboration ensuring AI complements rather than undermines human decision-making via counterfactual harm and complementarity principles.
Fine-tuned DeBERTaV3 system for lateral thinking in language models using humor/riddle data on BRAINTEASER task. LLM reasoning research.
Error correcting code based watermarking framework for detecting machine-generated text in language models. LLM safety research.
Game theory research on controlling strongly monotone games via gradient play and generalized Nash equilibrium constraints.
Quantum and classical neural networks for single-pixel imaging classification. Quantum ML research with imaging application.
Theoretical analysis of Temporal Difference learning convergence with linear function approximation. Foundational RL research.
Mamba-based Mixture of Experts architecture for EMG gesture recognition. Machine learning research with practical HCI application.
Evaluation of seven LLMs on health-claim verification across languages and contexts, assessing how linguistic and contextual factors affect accuracy of AI-generated health advice.
HoloLLM is a multimodal LLM for embodied agents in smart homes that processes diverse sensory inputs beyond vision for language-grounded perception and human behavior understanding.
Analysis showing watermarking degrades LLM alignment safety properties and proposes mitigation strategies for deployment compatibility.
Method enabling LLMs to generate and iteratively refine continuous control policies for embodied agent sensory-motor control.
Training method enabling LLMs to learn procedural knowledge from declarative instructions, demonstrating instruction efficiency in fine-tuning.
Comparative analysis of State Space Model and hybrid architectures versus Transformers for long-context processing on edge devices.
LLM-based agent system for schema-guided extraction and recommendation from financial tables with missing structural metadata.
Policy optimization method addressing credit assignment problems in reinforcement learning for aligning text-to-image generation models.
Framework for automatically generating demonstrations for training multi-step bimanual mobile manipulation robots under soft and hard constraints.
Research evaluating security vulnerabilities of backbone LLMs used in AI agents, addressing systematic security modeling for deployed agent systems.
Research on expert-router coupling loss for mixture-of-experts models to align router decisions with expert capabilities.
Fast-ThinkAct framework for efficient vision-language-action reasoning using compact latent planning to reduce inference latency.
Theoretical proof of universality for many-body quantum machine learning models in approximating quantum distributions.
FROST method for efficient LLM reasoning by pruning uncritical paths using attention weights to reduce inference latency.
Self-supervised learning framework applying vision models to cryo-EM density maps for structural biology analysis.
Research on object tracking using 3D geometric reasoning and online model editing to handle occlusion and appearance variations.
Technical report on UI-Venus-1.5, a unified GUI agent for automating digital environment interactions with multiple model variants.
Study on synthetic query generation for dense retrieval showing quality-diversity tradeoffs across 31 datasets, benefits multi-hop reasoning.
Research evaluating AI safety datasets, finding they overrely on obvious triggering cues and lack real-world adversarial depth.
Research identifies prompt injection attacks targeting LLM agent skills feature. Studies vulnerability of agent supply chains.
Research paper on video reasoning capabilities in vision models, studying spatiotemporal understanding and scaling behavior.
Blog post on designing AI agent personalities through system prompts and dynamic behavioral modeling beyond generic responses.
DeepSeek-V3.2 inference optimization on GB300 hardware using FP4 quantization and tensor parallelism, achieving 7360 TGS throughput.
GPT-5.3-Codex: agentic coding model combining frontier coding performance with reasoning capabilities, 25% faster than prior version, handles long-running research and tool-use tasks.
System running multiple AI agents in parallel for structured research tasks, demonstrated with stock analysis using specialized agents.
AgentWallet: governance layer for autonomous AI agents with dead man's switch, spend controls, and remote termination capabilities.
High-throughput LLM inference and serving engine with memory efficiency, developed at UC Berkeley.
Open-source self-hostable backend platform with authentication, database, file storage, and unified API.
Terminal CLI tool tracking usage, rate limits, and token costs across Claude, Codex, and Gemini providers.
Context window management technique for LLM agents using double-buffering to avoid lossy summarization during context exhaustion.
CPU-based symbolic reasoning engine achieving 18.1% on ARC-AGI-2 benchmark without neural networks or LLMs.
OpenPencil: open-source AI-native vector design tool with design-as-code philosophy and AI screen generation.
RLM-Codelens: codebase intelligence tool using recursive language models for repository analysis and architecture understanding.
Crustdata: web search API for AI agents with token-efficient design and entity mapping capabilities.
Claude Code configuration system for academic research workflows covering ideation, analysis, and publication phases.