Applied AI engineering

We build AI agents and run them in production.

Nina Labs designs agents, agentic workflows and AI software factories, connects them to the systems you already run, and measures them with evals. Everything we build lives in your repositories and your cloud.

We build with every major lab

Inside a run

What one agent run looks like

An accounts-payable agent processes a vendor invoice. It reads, checks, asks a person when the policy says so, and records every step.

Scoped tools
The agent can read the ERP and post invoices. It cannot pay them.
Policy in code
The 2% variance limit is a rule in code, not a line in a prompt.
Checks on every run
Six checks score each run before it is marked done.
ap-agent / run 7f3a21
Example run
model claude-sonnet-5-5 trigger inbox
  1. 00.0s input email · billing@kessler-logistics.example result: Invoice 4471
  2. 00.4s extract invoice-4471.pdf result: 14 lines · $18,240.00
  3. 02.1s tool erp.find_vendor("Kessler Logistics") result: V-0193
  4. 02.6s tool erp.match_po("PO-88210") result: 13 of 14 lines match
  5. 03.0s policy price variance 2.4% · limit 2.0% result: needs approval
  6. 03.1s approval sent to AP lead · #ap-approvals result: waiting
  7. 71.4s approved AP lead · "freight surcharge agreed"
  8. 71.9s tool erp.post_invoice(V-0193, PO-88210) result: batch 2210
  9. 72.3s done evals 6/6 passed · 9 tool calls result: $0.031
Example data, not from a client

01 Services

What we build and run

Eight service lines. Each one ends with a system in production, the evals that show it works, and a team on your side that can run it.

All services
  • 01

    AI agents in production

    We design, build and run AI agents that work inside your systems. They read, decide, call tools and hand the case to a person when a rule says so.

  • 02

    Agentic workflow automation

    We rebuild back-office processes as workflows. Fixed steps stay as code, agents handle the steps that need judgment, and people approve the exceptions.

  • 03

    AI software factory

    We set up coding agents across your delivery pipeline, from spec to merge. Your engineers own the specs, the review gates and the releases.

  • 04

    Claude and OpenAI team enablement

    We roll out Claude, ChatGPT, Claude Code and Codex to your teams, then teach each team to use them in its daily work, with admin setup, data rules and a library of tested workflows.

  • 05

    Legacy code modernization

    We use coding agents to document, test and migrate legacy systems in small, verified steps, from COBOL and old Java to aging .NET and monoliths.

  • 06

    MCP and systems integration

    We connect agents to your systems through Model Context Protocol (MCP) servers and APIs, with real authentication, scoped permissions and audit logs.

  • 07

    Evals, security and AI governance

    We measure whether your AI systems work, test how they fail and document them for auditors, with evals, red-team tests, tracing and compliance mapping.

  • 08

    AI strategy and operating model

    We help leadership pick the AI work that pays off, decide what to build or buy, and set up the operating model that keeps AI systems owned and measured.

02 Anthropic and OpenAI

Claude and ChatGPT, rolled out team by team

We roll out Claude, ChatGPT, Claude Code and Codex, connect them to your tools, and train each team on its own work. When a workflow is ready, we turn it into a production agent.

Team enablement

Anthropic

We work with
  • Claude
  • Claude Code
  • Claude Agent SDK
  • Claude Managed Agents

We roll out Claude to whole companies and build production agents on the Claude Agent SDK. Engineering teams get Claude Code set up on their own repositories.

  • Claude Enterprise setup: SSO, roles, data controls and connectors to your tools
  • Projects and Agent Skills for each team's repeat tasks
  • Claude Code for engineers: CLAUDE.md files, hooks, permissions, review rules and CI
  • Production agents on the Claude Agent SDK or Claude Managed Agents
  • Claude through the Claude API, Amazon Bedrock, Google Cloud or Microsoft Foundry

OpenAI

We work with
  • ChatGPT
  • Codex
  • Agents SDK
  • Responses API

We roll out ChatGPT Enterprise and Codex, train each team on its real work, and build agents on the OpenAI Agents SDK and the Responses API.

  • ChatGPT Enterprise setup: SSO, workspace roles, data controls and connectors
  • Shared projects and workflows for each team's repeat tasks
  • Codex for engineers: cloud tasks, CLI, IDE extension, AGENTS.md files and code review
  • Production agents on the OpenAI Agents SDK and the Responses API, including realtime voice
  • OpenAI models through the OpenAI API, Microsoft Foundry or Amazon Bedrock
  • Gemini Enterprise
  • Vertex AI
  • Microsoft Foundry
  • Amazon Bedrock
  • GitHub Copilot
Every lab and platform we use

03 Software factory

How an AI software factory works

Coding agents do most of the implementation work. Your engineers write and approve the specs, own the review gates and decide what ships.

  1. 01

    Ticket

    Engineer

    Describes the change and the acceptance criteria.

  2. 02

    Spec

    Agent → engineer

    An agent drafts the spec and the plan. An engineer approves them.

  3. 03

    Implement

    Coding agent

    Changes the code in a sandbox and runs the tests.

  4. 04

    Review

    Review agent

    Checks the change against the spec, your conventions and your security rules.

  5. 05

    Verify

    CI

    Runs the tests, the evals and the static analysis.

  6. 06

    Merge

    Engineer

    Reviews and merges. Your release process ships it.

Engineers keep control

Agents open pull requests. People approve specs, merge changes and decide what ships.

Agents run in sandboxes

Short-lived, least-privilege tokens. No production secrets in the context. Every run logged.

Measured from week one

Lead time, change failure rate, review time and the share of agent changes merged without rework.

How we set up a software factory

04 Approach

How an engagement runs

Four phases, with a working system in your environment by the end of the second. You see results on real cases before anything goes live.

  1. 01 1–2 weeks

    Map

    We sit with the people who do the work, measure the current process and pick the tasks where an AI system can finish most cases end to end.

    • Ranked use cases
    • Baseline metrics
    • Architecture and risk notes
  2. 02 4–8 weeks

    Build

    We build the system in your environment: the agent or workflow, the tools, the permissions and the eval suite. Your engineers pair with us from day one.

    • Working system in your cloud
    • Eval suite from real cases
    • Tracing and dashboards
  3. 03 2–4 weeks

    Prove

    The system runs in shadow mode on live cases while people do the real work. We compare the results and fix what the evals and the traces show.

    • Shadow-mode results
    • Go-live criteria
    • Runbook and on-call guide
  4. 04 Ongoing

    Run

    We turn it on in stages, watch quality and cost, and re-run the evals on every model or prompt change. Then we hand it over to your team.

    • Staged rollout
    • Monthly quality and cost report
    • Handover and exit plan

05 Standards

What every system we ship includes

These are part of the build, not extras. If one is missing, the system does not go live.

01 Evals
A suite of real cases runs on every change to prompts, tools or models.
02 Tracing
Every step, model call and tool call is recorded in OpenTelemetry format.
03 Approvals
Actions that move money, change records or contact customers need a person until the data says otherwise.
04 Permissions
Each tool has a written scope. Agents act with the user's rights or a narrow service identity.
05 Budgets
Cost, token and latency limits per run, with alerts before they are reached.
06 Hosting
Your cloud account or VPC by default. Open-weight models when data must stay on your hardware.
07 Models
Chosen per task on measured quality, latency and cost. Switching is a configuration change.
08 Ownership
Code, prompts, eval sets and infrastructure live in your repositories from the first day.
09 Handover
A runbook, training for your team and a written exit plan in every engagement.

06 Stack

We build on the platforms you already use

Every major lab, cloud and open-weight model. We do not resell any of them: we pick per task, build on open protocols and keep your option to switch.

All platforms

Frontier labs

  • Anthropic Claude
  • OpenAI GPT
  • Google Gemini
  • Meta
  • Mistral AI
  • Perplexity
  • Cohere
  • Grok

Open-weight models

  • DeepSeek
  • Z.ai GLM
  • Qwen
  • Kimi
  • MiniMax
  • Llama
  • Gemma
  • NVIDIA Nemotron

Agent platforms

  • Claude Agent SDK
  • OpenAI Agents SDK
  • Google ADK
  • Vertex AI
  • Microsoft Foundry
  • Amazon Bedrock
  • LangGraph
  • LlamaIndex

Coding agents

  • Claude Code
  • OpenAI Codex
  • GitHub Copilot
  • Cursor
  • Kiro
  • Google Antigravity
  • Windsurf

Protocols

  • Model Context Protocol
  • Agent2Agent (A2A)
  • AGENTS.md
  • Agent Skills

Clouds and runtimes

  • AWS
  • Microsoft Azure
  • Google Cloud
  • NVIDIA NIM
  • Hugging Face
  • vLLM
  • Ollama
  • Your own hardware

Orchestration

  • Temporal
  • n8n
  • AWS Step Functions
  • Azure Durable Functions

Observability and evals

  • OpenTelemetry
  • Langfuse
  • LangSmith
  • Braintrust
  • Arize Phoenix
  • Promptfoo

07 Lab

We build our own products too

We build and run our own AI products on the same stack, and to the same standards, as our client work.

See the lab
git1file Product

Code repositories, packed for LLMs

Turns a GitHub repository into one compact, structured file for a model's context, with at least 25% fewer tokens than other tools. Launched in March 2025.

Visit git1file.com
Microsoft for Startups Program

Member since January 2025

Azure credits, technical guidance and access to Microsoft's startup network for the products and client systems we build.

Read the announcement

08 FAQ

Questions we hear often

What does Nina Labs do?

We are an applied AI engineering firm. We design, build and run AI agents, agentic workflows and AI software factories inside companies, connect them to the systems those companies already use, and measure them with evals.

What is an AI software factory?

A delivery pipeline where coding agents do most of the implementation work. Engineers write and approve the specs, own the review gates and decide what ships. We set up the agents, the repository groundwork, the review and eval gates, and the metrics.

Which models and platforms do you work with?

We work with every major lab: Anthropic (Claude), OpenAI (GPT, ChatGPT and Codex), Google (Gemini and Vertex AI), Meta, Mistral and Perplexity, and with open-weight models from DeepSeek, Z.ai, Qwen, Kimi and MiniMax. We run them on Amazon Bedrock, Microsoft Foundry, Google Cloud or your own hardware, pick the model per task on measured quality, latency and cost, and build on open protocols such as MCP and A2A so you can switch later.

Where do the systems run, and who owns them?

In your cloud account by default. The code, prompts, eval sets and infrastructure definitions live in your repositories from the first day. You own all of it.

How do you control what an agent can do?

Every tool has a written scope, agents act with the user's permissions or a narrow service identity, and actions that move money, change records or contact customers need a person's approval until the evals and production data show otherwise. Every run is traced.

How long does a first project take?

The mapping phase takes one to two weeks. A first system usually reaches shadow mode four to eight weeks after that, and goes live in stages over the following weeks.

Can you train our teams to use Claude and ChatGPT?

Yes. We set up Claude Enterprise, ChatGPT Enterprise or both, write the data rules with your security team, and run hands-on workshops for each team on its own tasks. Engineers get Claude Code and Codex set up on their repositories. The result is a library of tested workflows that each team owns.

How do we start?

Send us a short note through the contact page. An engineer replies, usually with a few questions and times for a 30-minute call. If an AI system is not the right tool for the problem, we will say so.

Tell us which process you want to hand to an agent

A 30-minute call with an engineer. We will tell you whether an AI system is the right tool for it, and what it would take to run it in production.