AI Platform · Robotics Cloud · Distributed Systems

AI Platform Architecture Portfolio

A focused collection of personal engineering labs and open-source work for AI platform, robotics cloud and distributed systems roles.

Production experience is described only at a verified and sanitized level. Experiments are not presented as production scale. Articles may be AI-assisted; code, tests and reproductions are verified by me.

Selected Engineering Work

01 · Robotics Cloud to AI Platform

Robotics cloud work across OpenAPI, device state, task scheduling, asynchronous processing, idempotency, recovery and production observability. The platform mapping for companion robots focuses on identity, session state, tool calls, model routing, long-term memory and failure recovery.

Discuss: tenant isolation, device/family-member boundaries, reconnects, permissions and high-risk action confirmation.

Proxmox VE + AI Agent →

02 · AI Video Studio

A provider-neutral AI content pipeline covering Planner, Shot, Queue, Worker, Provider, Storage and Render.

Discuss: idempotency, worker leases, retries, provider throttling, async TTS/video jobs, cost accounting, MCP boundaries and observability.

Source and tests →

03 · Multi-Agent Shared Memory

Experiments around shared context, memory extraction, storage, retrieval, conflict handling and isolation across agent frameworks.

Discuss: freshness, correction, family-member isolation, sensitive-field masking, degradation and rollback.

Shared memory practice →

04 · AI Gateway / Quota / Cost

A self-hosted dashboard and provider adapters for model quota, cost, request state and failure analysis.

Boundary: authorized credentials only; no cookie storage, service-limit bypass or third-party session collection; secrets come from environment variables or a Secret Store.

Dashboard practice →

05 · Observability and Recovery

Logs, metrics, traces, event timelines, capacity and incident review connected into an evidence chain rather than a single “service is healthy” screenshot.

Discuss: P95/P99 latency, retry storms, circuit breaking, queue buildup, worker recovery, alert noise and capacity planning.

Observability practice →

How I Work