Applied AI — Freight Tender Automation
Our Customer Service team at STG Logistics was manually reading and keying up to 200,000 freight tenders a year out of email — up to 16,700 labor hours annually just on data entry. I identified the bottleneck, ruled out the obvious fix, designed an LLM-based intake pipeline to replace it, and shipped a system that now processes 95% of tenders touchless.
What I Found
Freight tenders reach us one of two ways. EDI (204/990 or equivalent) is structured, machine-readable, and already integrated into our TMS — it runs with effectively no human touch. Email is unstructured, customer-specific, and entirely dependent on a person reading it correctly. Email wasn't a shrinking edge case for us — it was the majority of daily volume, because a large share of our shippers, especially mid-market and regional accounts, don't have the EDI capability (or the will) to integrate directly.
The difficulty was never volume — it was variance. Across customers, a “tender” could show up as freeform text in the email body, a PDF attachment (sometimes scanned, sometimes a “printed” digital document), an Excel/CSV attachment with a customer-specific column layout, or an attachment forwarded within another email so the real tender sat a layer deep — sometimes several of those in the same message, and not always in agreement. Each format could carry 50–100 fields per order: origin/destination, appointment windows, commodity, weight, equipment type, accessorials, reference numbers, rate details. There was no single schema; every customer effectively had their own dialect, and it could change without notice. A rep had to identify what kind of message it was, extract the right fields into the right TMS fields, resolve ambiguity, and key it in — order by order. That's a judgment-heavy task disguised as data entry, which is exactly why no one had automated it with simple tools, and exactly why it was expensive to run manually at 200K events a year.
What I Ruled Out
The instinctive proposal — “run OCR on the attachments” — sounded reasonable and failed for three structural reasons.
OCR solves “read the characters.” The problem I actually had was “understand an unfamiliar document and produce a structured order,” which is a language-and-reasoning problem, not a vision problem.
What I Designed
Large language models changed the economics here because they don't require a template per customer — I could give the model a target schema (our TMS order fields) and a document, and have it extract and normalize data against that schema regardless of layout. The model reasons over context the way a trained rep would, it's multi-modal out of the box (email body, PDF, scanned images, Excel), it can classify where a tender actually lives before extracting — body, attachment, or forwarded attachment — and it can flag fields it's unsure about instead of silently guessing, which is what makes a human-in-the-loop design workable instead of risky.
I didn't design this to let the model write directly to the TMS unsupervised. Here's the shift, side by side:
The new pipeline forwards each inbound email to an AI classifier that decides whether it's actually a tender. If it is, the model reads and validates the data, and on success maps the fields directly to an EDI 204 so the TMS can process it like any other structured tender and create the order automatically. If extraction fails or mandatory fields are missing, the system logs it and notifies the rep — who now only touches the tenders that need a human, instead of every single one. That's the meaningful shift: reps stopped being data-entry clerks and became exception handlers.
What I Built
The “AI Layer” in that diagram isn't a black box — it's a chain of narrow, single-purpose reasoning steps rather than one giant prompt trying to read, extract, validate and format in a single pass. An inbound email triggers the chain; a sequence of GPT-5.5 reasoning steps classify the message and extract and normalize the tender fields; a validation gate (length(order.errors)==0) checks the result against mandatory-field rules; and only a clean tender reaches the function call that generates the EDI 204 and hands it to the TMS. Anything that fails the gate skips the EDI step and routes back to a rep instead.
Keeping each step narrow and single-purpose is what makes the confidence scoring trustworthy — a step that only ever does one job is far easier to grade, and far easier to catch failing, than one model call trying to do everything at once.
What I Delivered
Once the pipeline was live and tuned, the numbers moved fast:
| Metric | Before | After |
|---|---|---|
| Tender processing | Fully manual, tender-by-tender | 95% touchless |
| Order-entry effort | Baseline | 80% reduction |
| CS Rep time | 100% data entry | 60% redirected to customer-facing work |
| Order routing | Manual TMS entry | Auto-mapped to EDI 204, exceptions routed to a rep |
This wasn't just cost-out. It was capacity: the same team could absorb volume growth without proportional hiring, and the time we got back went straight into the judgment calls — carrier sourcing, exception resolution, customer escalations — that actually need a person.
What I'd Do Again
Closing Thought
The instinctive fix — throw OCR at unstructured documents — treats this as a character-recognition problem. It isn't. It's a document-understanding problem wearing a data-entry costume, and it responds to tools built for language and reasoning, not tools built for reading pixels off a fixed template. I didn't just automate a workflow — I changed what my Customer Service team is actually for.
Concepts & methods