Goolean

AI & LLM

Production AI for work that can't afford to be wrong.

Goolean builds and operates its own large language model, Mike, purpose-built for document-heavy, regulated workflows and running entirely on our appliance inside the client's environment. We bring that same engineering — evaluation harnesses, policy engines, evidence ledgers, mutation-tested guards — to AI products for our clients.

What's included

Capabilities

01

Mike — Goolean's own LLM

A domain-specialised model stack for legal and financial document work, delivered pre-installed on an appliance in your building. Inference runs there; Mike Cloud only pushes signed releases and knowledge packs over an outbound-only channel. No document or PII ever leaves your network.

02

Fine-tuning & adaptation

Per-client LoRA adapters trained on licensed or synthetic data, with a partition firewall between training, development and acceptance sets.

03

Document intelligence

OCR, classification, key-field extraction, PII redaction and cross-document consistency — with a deterministic safety net under every model.

04

Evaluation & safety harnesses

Gold datasets, scorers, release gates and a Pareto frontier of cost vs quality for every model version. Guards are mutation-tested.

05

Governed runtimes

Policy-as-data engines (default deny), typed task contracts, append-only hash-chained ledgers. The model observes; policy decides.

06

LLM application development

Retrieval, agents, structured output and human-in-the-loop review, built on whichever model — ours or a frontier API — the harness says wins.

How we work

From first call to steady state.

  1. Step 1

    Define the expensive failure

    Before models: what is the wrong answer that costs you money or a licence? That becomes the first test, and the release gate.

  2. Step 2

    Build the harness

    Gold data, scorers and a fake backend. Everything above the model is green before real inference runs.

  3. Step 3

    Bake off

    Candidate models — Mike, open weights, frontier APIs — compete on your data under identical rules. Licence and residency are enforced in code.

  4. Step 4

    Deploy where the data lives

    On-prem appliance, VPC or edge node. Updates arrive as signed releases over an outbound-only channel, pass on-box acceptance before activation, and roll back by manifest. Documents and PII never travel.

What you get
  • ✓Raw data local-only by construction, enforced by tests, not policy documents
  • ✓Every release gated on a frozen acceptance set, evaluated once
  • ✓No model-generated number ever enters a filing or a report
  • ✓An auditable ledger of what ran, on which model, with what result
Typical stack
MikeApple MLXPyTorchQwen-VLLoRAvLLMTesseractpdfplumberPyMuPDFPostgreSQLFastAPIpytestoutlines

Case studies

Related work

Anonymised where the client asked for it. Every outcome listed is one we can substantiate.

Questions

Frequently asked

Why your own LLM rather than an API?+

Because for regulated document work the data cannot leave the building, the outputs must be reproducible, and the vendor must be accountable for the model's behaviour. Mike gives us all three; APIs give none.

Can you still use OpenAI, Anthropic or Google models?+

Yes, where the data classification allows it. Our harness treats every model as a replaceable component and picks the winner on evidence.

What hardware does Mike need?+

The reference appliance is a Mac with 128 GB unified memory; NVIDIA deployments are supported for higher throughput. We size it to your page volume.

Other services

Talk to us about ai & llm engineering.

One scoping call with a delivery lead and an architect. We'll say plainly whether — and how — we'd take it on.

Talk to us