Cabinly Technologies · Engineering

Engineering

Production AI systems, agentic workflows and the discipline that keeps them reliable.

5 capability areas · 40+ agent skills in daily use · as of 27 Sep 2026

Reference architecture

A generative pipeline with a QA gate.

Six stages and an automated quality loop. Failed shots are regenerated before assembly.

Production pipeline: Idea, Script, Characters, Scenes, Voice and captions, Masters and Shorts, then a QA loop that sends failed shots back to Scenes.01Idea02Script03Characters04Scenes05Voice &captions06Masters &Shorts07QA loopfailed shots go backProduction pipeline: Idea, Script, Characters, Scenes, Voice and captions, Masters and Shorts, then a QA loop that sends failed shots back to Scenes.01Idea02Script03Characters04Scenes05Voice & captions06Masters & Shorts07QA loopfailed shots go back

Capabilities

What we build and run in production.

What it does, the stack, and the detail one tap away.

  1. Agentic engineering

    Spec-driven agent workflows that plan, build, review and verify software, with a human approving each stage.

    • Claude Code
    • Custom agent skills
    • Subagent orchestration
    • Playwright
    • MCP integrations
    • Specification → plan → task-scoped agents → independent review → verification
    • 40+ versioned agent skills, each a written, reviewable playbook
    • Every agent task ends in tests and a separate reviewer pass, not self-assessment
    • Browser-level verification of real screens before work is accepted
    • Shared context layer: a knowledge hub mapping projects, dependencies and ownership
    • Guardrails: scoped permissions, no destructive actions without approval, audit trail per task

    In detail

    We use agents the way a disciplined engineering team uses people.

    We use agents the way a disciplined engineering team uses people. Work starts as a written specification. It becomes a plan of small tasks with exact files, interfaces and tests. Each task goes to an agent scoped to that task alone. A separate reviewer agent checks the result against the specification before anything moves on.

    Humans decide. Agents propose, implement and verify. Specifications, plans, review findings and decisions are all written down, so every change can be traced back to why it was made.

    Our library of more than 40 agent skills covers recurring work end to end: building and testing software, running production pipelines, publishing, and research. The same approach is available to clients who want agentic workflows built into their own teams.

  2. LLM product backends

    Production conversational AI on serverless infrastructure, with streaming, long-term memory and cost-aware model routing.

    • Python FastAPI
    • AWS Lambda (streaming)
    • PostgreSQL
    • Redis
    • Anthropic and OpenAI SDKs
    • Sentry
    • Offline evals
    • Streaming responses from FastAPI on AWS Lambda
    • Long-term memory with recall and supersession, so facts stay current
    • Runtime-configurable routing across multiple LLM providers, with prompt caching
    • One backend serving two live products; every change is checked against both
    • Offline evaluations before prompt or model changes ship
    • Error monitoring and managed PostgreSQL and Redis

    In detail

    Two live consumer products, Cabin and Mo (the AI companion inside QuesMo), run on one shared backend.

    Two live consumer products, Cabin and Mo (the AI companion inside QuesMo), run on one shared backend. It handles streaming chat, long-term memory, voice and media, and subscriptions, and it scales with demand on serverless infrastructure.

    Memory is designed for accuracy over volume. It recalls what matters and replaces facts that have changed, instead of accumulating contradictions. Model routing is configurable at runtime, so cost and quality can be balanced per request without a redeploy. Prompt caching keeps repeated context cheap.

    Changes to prompts or models go through offline evaluations first. Because two products depend on the same service, every change is verified against both before it ships.

  3. Automation pipelines

    Scheduled, event-driven pipelines that turn raw inputs into reviewed, publish-ready content.

    • Python
    • AWS Lambda (container images)
    • EventBridge Scheduler
    • DynamoDB
    • Pillow
    • Playwright
    • Claude
    • Scheduled ingestion of trending topics, twice a day
    • LLM drafting into a review queue, never straight to production
    • Paced publishing so output is spread across the day
    • Deduplication and validation before anything goes live
    • Server-side image rendering for branded quote cards

    In detail

    A consumer app needs fresh content every day, and writing it all by hand does not scale.

    A consumer app needs fresh content every day, and writing it all by hand does not scale. A scheduled serverless pipeline gathers trending topics, drafts questions and posts with an LLM, and places them in a queue for review.

    A separate publisher releases approved items at a steady pace through the day. Deduplication and validation run before anything is published.

    The same pattern is being extended to an AI character that posts in its own voice: generate, render an image or quote card, check for duplicates, then publish.

  4. Generative media pipeline

    A shared production pipeline from script to finished episode, with automated visual QA before assembly.

    • Python
    • FFmpeg
    • Playwright automation
    • Generative video, image and voice models
    • Automated visual QA
    • Script and research first, with sources kept alongside each episode
    • Locked character references keep characters consistent across scenes and seasons
    • Image-to-video, voice and shot-by-shot assembly from one shared toolchain
    • Hindi and English caption rendering
    • One 16:9 master produces 9:16 cut-downs and thumbnails automatically
    • Automated visual QA scores sampled frames and regenerates failed shots

    In detail

    Every studio production runs through one shared pipeline.

    Every studio production runs through one shared pipeline. It starts with a script and research, then locks each character to reference images so they stay consistent from the first scene to the last. Scenes are generated, assembled shot by shot against the script, and captioned.

    The pipeline produces a 16:9 master, then derives vertical cut-downs and thumbnails from it. Because it is shared, an improvement made once applies everywhere.

    Quality is checked by machine before a person reviews it. A QA loop samples frames from every clip, scores them against a per-production checklist, and regenerates any shot that fails before assembly.

  5. Research and evaluation

    Recurring, ranked technical research that informs model, tooling and product decisions.

    • Claude Code
    • Web research
    • arXiv
    • GitHub
    • Hugging Face
    • Hacker News
    • Scheduled digests on models, agent tooling, LLM cost and quality, and media tools
    • Findings ranked against our own products and priorities
    • The top finding in each cycle researched in depth and archived with a date
    • Multi-source research combining papers, repositories, and web and video sources
    • Idea testing with simulated audience panels, always labelled as simulated

    In detail

    Scheduled research routines scan papers, open-source releases and technical communities.

    Scheduled research routines scan papers, open-source releases and technical communities. Each finding is ranked by its relevance to our products. The most important one is researched in depth and archived as a dated digest.

    This feeds concrete decisions: which models to route to, which tools to adopt, and where cost or quality can be improved.

    New ideas are pressure-tested before they are built. Simulated audience panels react to each concept, and results are always labelled as simulated, never presented as real customer research.

How we build

Specify, plan, build, verify.

One loop for our products and client work. Every step is written down.

  1. Specify

    The outcome, users, constraints and acceptance criteria, written down.

  2. Plan

    Small, testable tasks, each with exact files, interfaces and tests.

  3. Build and review

    A scoped agent builds each task test-first. A separate reviewer checks it against the spec.

  4. Verify

    Tests, accessibility and privacy checks, and a real browser before release.

A failed check sends the task back to step 03.

Engineering standards

What holds on every project.

Work with us

Need agentic AI built properly?

LLM products, agent workflows and automation pipelines, from spec to production.

Start a conversation