← Back to case studies

Applied AI — Freight Tender Automation

From Inbox to Order: Automating Tender Processing at Scale

Our Customer Service team at STG Logistics was manually reading and keying up to 200,000 freight tenders a year out of email — up to 16,700 labor hours annually just on data entry. I identified the bottleneck, ruled out the obvious fix, designed an LLM-based intake pipeline to replace it, and shipped a system that now processes 95% of tenders touchless.

CompanySTG Logistics Inc.
RoleVP, Product Management & BI/Analytics
Timeline2020–2026
CategoryApplied AI / Intelligent Document Processing
95%Touchless tender processing
80%Reduction in order-entry effort
60%Team time redirected to customers

What I Found

Two channels, one bottleneck

Freight tenders reach us one of two ways. EDI (204/990 or equivalent) is structured, machine-readable, and already integrated into our TMS — it runs with effectively no human touch. Email is unstructured, customer-specific, and entirely dependent on a person reading it correctly. Email wasn't a shrinking edge case for us — it was the majority of daily volume, because a large share of our shippers, especially mid-market and regional accounts, don't have the EDI capability (or the will) to integrate directly.

A real inbound freight tender received by email, with contact and reference fields redacted
A representative inbound tender — ~30 fields across shipper, consignee, references and charges, with no fixed schema and no two customers formatted the same way. (Contact and reference fields redacted for privacy.)

The difficulty was never volume — it was variance. Across customers, a “tender” could show up as freeform text in the email body, a PDF attachment (sometimes scanned, sometimes a “printed” digital document), an Excel/CSV attachment with a customer-specific column layout, or an attachment forwarded within another email so the real tender sat a layer deep — sometimes several of those in the same message, and not always in agreement. Each format could carry 50–100 fields per order: origin/destination, appointment windows, commodity, weight, equipment type, accessorials, reference numbers, rate details. There was no single schema; every customer effectively had their own dialect, and it could change without notice. A rep had to identify what kind of message it was, extract the right fields into the right TMS fields, resolve ambiguity, and key it in — order by order. That's a judgment-heavy task disguised as data entry, which is exactly why no one had automated it with simple tools, and exactly why it was expensive to run manually at 200K events a year.

What I Ruled Out

Why traditional OCR wasn't the answer

The instinctive proposal — “run OCR on the attachments” — sounded reasonable and failed for three structural reasons.

  • OCR reads characters, not meaning. Classic OCR converts pixels to text. It doesn't know that the string sitting 40 pixels below “Ship To” is a destination address versus a bill-to address versus a warehouse note — without a fixed template, it has no way to attach meaning to position.
  • Template-based extraction breaks on variance. Mapping fields per customer by coordinate or anchor text works until a customer tweaks their PDF export or sends a one-off variant. Across our customer base, template maintenance would have become its own full-time job — and it still wouldn't touch email-body tenders or nested attachments at all.
  • It doesn't generalize across modality. OCR addresses images and scanned PDFs. It does nothing for tenders sitting in plain email text or in an Excel sheet with a bespoke column layout — solving for “PDF only” would have left most of our actual variance untouched.

OCR solves “read the characters.” The problem I actually had was “understand an unfamiliar document and produce a structured order,” which is a language-and-reasoning problem, not a vision problem.

What I Designed

An LLM-based intake pipeline, not another template engine

Large language models changed the economics here because they don't require a template per customer — I could give the model a target schema (our TMS order fields) and a document, and have it extract and normalize data against that schema regardless of layout. The model reasons over context the way a trained rep would, it's multi-modal out of the box (email body, PDF, scanned images, Excel), it can classify where a tender actually lives before extracting — body, attachment, or forwarded attachment — and it can flag fields it's unsure about instead of silently guessing, which is what makes a human-in-the-loop design workable instead of risky.

I didn't design this to let the model write directly to the TMS unsupervised. Here's the shift, side by side:

Before AI: a single-lane manual workflow where the CS Rep reads, validates and keys every tender by hand
Before — every tender is read, validated and keyed into the TMS by hand, one rep, one lane, start to finish.
After AI: a four-lane workflow across Email Client, AI Layer, TMS and CS Rep, where the AI classifies, extracts, validates and maps tenders to EDI 204, only escalating exceptions
After — the AI Layer classifies, extracts, validates and maps straight to EDI 204; the CS Rep only sees what fails validation or is missing mandatory fields.

The new pipeline forwards each inbound email to an AI classifier that decides whether it's actually a tender. If it is, the model reads and validates the data, and on success maps the fields directly to an EDI 204 so the TMS can process it like any other structured tender and create the order automatically. If extraction fails or mandatory fields are missing, the system logs it and notifies the rep — who now only touches the tenders that need a human, instead of every single one. That's the meaningful shift: reps stopped being data-entry clerks and became exception handlers.

What I Built

Under the hood: the actual reasoning chain

The “AI Layer” in that diagram isn't a black box — it's a chain of narrow, single-purpose reasoning steps rather than one giant prompt trying to read, extract, validate and format in a single pass. An inbound email triggers the chain; a sequence of GPT-5.5 reasoning steps classify the message and extract and normalize the tender fields; a validation gate (length(order.errors)==0) checks the result against mandatory-field rules; and only a clean tender reaches the function call that generates the EDI 204 and hands it to the TMS. Anything that fails the gate skips the EDI step and routes back to a rep instead.

The live LLM orchestration graph: an email trigger feeding a sequence of GPT-5.5 reasoning steps, a mandatory-field validation gate, and a function call that sends the EDI 204
The live orchestration graph — trigger, chained GPT-5.5 reasoning steps, a validation gate, and the EDI 204 hand-off to the TMS.

Keeping each step narrow and single-purpose is what makes the confidence scoring trustworthy — a step that only ever does one job is far easier to grade, and far easier to catch failing, than one model call trying to do everything at once.

What I Delivered

The results

Once the pipeline was live and tuned, the numbers moved fast:

MetricBeforeAfter
Tender processingFully manual, tender-by-tender95% touchless
Order-entry effortBaseline80% reduction
CS Rep time100% data entry60% redirected to customer-facing work
Order routingManual TMS entryAuto-mapped to EDI 204, exceptions routed to a rep
Live operations dashboard tracking tender volume, emails processed, and automated orders, with customer-level detail redacted for privacy
The live operations dashboard I use to track this in production — tender volume, touchless automation rate, and daily throughput. (Customer-level detail redacted for privacy.)

This wasn't just cost-out. It was capacity: the same team could absorb volume growth without proportional hiring, and the time we got back went straight into the judgment calls — carrier sourcing, exception resolution, customer escalations — that actually need a person.

What I'd Do Again

Implementation lessons

  • I started with our highest-volume, most-templated customers, not the long tail of one-off formats — proving the confidence-and-validation loop where ROI was largest and variance was most manageable, then expanding.
  • I made validation against TMS reference data non-negotiable. Extraction without a check against known locations, rate agreements, and equipment codes is how a bad tender becomes a bad order.
  • I tracked straight-through-processing rate as the north star, not raw extraction accuracy — a model can be 95% accurate on fields and still be unsafe to auto-file if the 5% it misses are the fields that matter.
  • I treated the human review queue as a data asset. Every correction a rep makes is a labeled example, so the system keeps improving instead of sitting as a static deployment.
  • I kept governance in the loop. Customer contracts, EDI trading partner agreements, and data handling requirements didn't disappear because the intake mechanism changed — the automation had to sit inside those constraints, not around them.

Closing Thought

It was never a character-recognition problem

The instinctive fix — throw OCR at unstructured documents — treats this as a character-recognition problem. It isn't. It's a document-understanding problem wearing a data-entry costume, and it responds to tools built for language and reasoning, not tools built for reading pixels off a fixed template. I didn't just automate a workflow — I changed what my Customer Service team is actually for.

Concepts & methods

Intelligent Document Processing LLM-based Extraction Vision-language Models Schema-driven Parsing EDI 204 Mapping Chained Reasoning Agents Confidence-based Routing Human-in-the-loop Review Reference-data Validation TMS Integration Process Automation Freight Tender Management