Air Waybill OCR and Sea Waybill OCR: Extract Shipping Document Data

August 19, 2026

Answer: Air waybill OCR and sea waybill OCR use AI document extraction to read shipping documents and turn them into structured data for freight, customs, accounting, and logistics workflows. Lido is the best fit when you need to process air waybills, MAWBs, HAWBs, sea waybills, ocean waybills, and other carrier-specific layouts without templates or model training. It extracts fields like AWB number, shipper, consignee, carrier, routing, vessel, voyage, container number, seal number, cargo description, weight, charges, and customs values, then exports the results to Excel, CSV, Google Sheets, JSON, API, or downstream workflow systems.

Waybills look standardized until you process them at volume. One airline uses one air waybill layout, a freight forwarder sends a house air waybill in a different format, an ocean carrier uses its own sea waybill template, and the supporting documents arrive as PDFs, scans, photos, and email attachments.

If your team is manually keying those documents into a freight management system, TMS, customs platform, spreadsheet, or accounting workflow, the problem is not just OCR. You need reliable shipping document data extraction: the document has to be classified, the right fields have to be pulled from the right place, and the output has to be consistent enough to import downstream.

What air waybill OCR and sea waybill OCR actually need to do

Basic OCR turns an image into text. That is not enough for air waybill data extraction or sea waybill data extraction. This is also why generic PDF-to-Excel converters fail on trade documents once layouts, attachments, and field definitions vary.

A production workflow needs to understand the document, not just read it. The software has to know that an AWB number is different from a booking number, that a carrier prefix is not a flight number, that a port of discharge is different from a place of delivery, and that a container number, seal number, and package count should not be mixed together.

The workflow usually looks like this:

  1. Ingest the documents. Air waybills, sea waybills, invoices, packing lists, bills of lading, and supporting files arrive by email, upload, shared drive, or API.
  2. Classify the document type. The system identifies whether the file is an air waybill, master air waybill, house air waybill, sea waybill, ocean bill of lading, commercial invoice, or another logistics document.
  3. Extract the required fields. The system pulls shipment identifiers, parties, routing, cargo details, weights, charges, references, and any custom fields your team needs.
  4. Normalize the data. Dates, airport codes, port names, weights, currencies, and reference numbers are formatted consistently.
  5. Review exceptions. Humans check fields that are missing, conflicting, or business-critical before the data is imported.
  6. Send the output downstream. Data goes to Excel, CSV, Google Sheets, JSON, an API, a customs system, freight forwarding software, a TMS, or an accounting workflow.

That is the difference between OCR that gives you text and waybill OCR software that removes manual data entry.

Why air and ocean waybills break generic OCR tools

Air waybills and sea waybills are document families, not single fixed templates. That is why template-based OCR often breaks in logistics operations.

Common problems include:

  • Carrier-specific formats. Airlines, freight forwarders, ocean carriers, and NVOCCs each use different layouts.
  • MAWB and HAWB differences. A master air waybill and house air waybill can describe related shipments at different levels of consolidation.
  • Similar-looking fields. AWB numbers, booking numbers, bill of lading numbers, shipment references, and customer references can appear close together.
  • Scans and email attachments. Many documents arrive as forwarded PDFs, compressed scans, phone photos, or documents embedded in long email threads.
  • Mixed document packets. A single file may include an air waybill, commercial invoice, packing list, certificate of origin, delivery order, and carrier invoice.
  • Multi-language and international formats. Ports, airports, consignee addresses, carrier names, and commodity descriptions often span countries and languages.
  • Downstream risk. A wrong weight, container number, seal number, customs value, or consignee can create billing, clearance, release, or tracking problems.

For low-volume work, a person can open the PDF and key the data manually. At scale, that becomes a bottleneck. The goal is not to remove every human from the process. The goal is to stop making people retype every field and only have them review the fields that need judgment.

Air waybill OCR: fields that matter

An air waybill, or AWB, is the core document for air freight. It is also called an air consignment note. It is typically a non-negotiable document that functions as a receipt, contract of carriage, tracking document, and source of shipment details for billing, insurance, and customs workflows.

Air waybill OCR software should handle both clean digital PDFs and scanned or photographed AWBs. It should also support the common spelling variations people use in search and operations: air waybill OCR, airway bill OCR, AWB OCR, AWB data extraction, MAWB OCR, and HAWB OCR. For a focused workflow example, see air waybill OCR software powered by Lido.

Typical air waybill fields include:

Field group Examples to extract
Shipment identifiers AWB number, MAWB number, HAWB number, carrier prefix, shipment reference, tracking number
Parties Shipper, consignor, consignee, account numbers, notify party, agent, freight forwarder
Air routing Origin airport, destination airport, routing, airline carrier, flight number, departure date, destination codes
Cargo details Number of pieces, package type, commodity description, HS code, dimensions, gross weight, chargeable weight
Value and charges Declared value, customs value, prepaid charges, collect charges, rate class, currency, insurance details
Instructions and compliance Handling instructions, dangerous goods notes, temperature requirements, special service codes, issue date, place of execution

The AWB number is especially important. A standard AWB number is usually an 11-digit identifier made up of a carrier prefix, a serial number, and a check digit. If that number is wrong, shipment tracking, billing, and exception handling can fail downstream.

Sea waybill OCR: fields that matter

A sea waybill is used for ocean freight. Like an air waybill, it is generally non-negotiable. It serves as evidence of the contract of carriage and receipt of goods, but it is not a document of title. That is one of the key differences between a sea waybill and a negotiable ocean bill of lading.

Sea waybill OCR is also searched as sea way bill OCR, seawaybill OCR, SWB OCR, ocean waybill OCR, sea waybill data extraction, and shipping document OCR. The underlying need is the same: extract structured shipment data from ocean transport documents without manually building a template for every carrier or NVOCC. For a focused workflow example, see sea waybill OCR software powered by Lido.

Typical sea waybill fields include:

Field group Examples to extract
Shipment identifiers Sea waybill number, booking number, carrier reference, customer reference, bill of lading reference
Parties Shipper, consignee, notify party, carrier, freight forwarder, NVOCC
Ocean routing Vessel, voyage, port of loading, port of discharge, place of receipt, place of delivery, final destination
Container details Container number, seal number, container type, number of packages, marks and numbers
Cargo details Commodity description, HS code, gross weight, net weight, measurement, CBM, package type
Charges and instructions Freight charges, prepaid or collect terms, special instructions, issue date, place of issue

The risky fields are usually the ones that drive release, matching, billing, or customs workflows: container numbers, seal numbers, ports, weights, package counts, consignee names, and shipment references. Those are the fields where automated extraction should be paired with validation rules and exception review.

AWB, MAWB, HAWB, sea waybill, and bill of lading are not interchangeable

One reason logistics OCR projects get messy is that teams use “waybill,” “bill of lading,” “AWB,” and “shipping document” loosely. Your software needs to separate them because the extracted fields and downstream workflows are different.

Document Mode What it usually represents OCR implications
Air waybill Air Non-negotiable air freight document used as receipt, contract, tracking, and shipment detail record Extract AWB number, airline, airport routing, pieces, weights, charges, shipper, consignee, and handling instructions
Master air waybill Air Carrier-issued air waybill for a consolidated shipment Useful for forwarder and consolidation workflows; may need matching to one or more HAWBs
House air waybill Air Forwarder-issued document for an individual shipment inside a consolidation Extract individual shipper, consignee, and cargo details; often matched to a MAWB
Sea waybill Ocean Non-negotiable ocean transport document and receipt of goods Extract carrier, vessel, voyage, ports, containers, seals, cargo, weights, and references
Ocean bill of lading Ocean May be negotiable and can function as a document of title depending on type Often needs additional title, release, endorsement, and compliance handling beyond simple OCR

If your team processes all of these documents, do not choose a tool that only works on one fixed form. Choose a logistics document OCR system that can classify the file first, then apply the right extraction schema for each document type. If bills of lading are a major part of the workflow, see our guide to the best bill of lading OCR software.

How Lido extracts air waybill and sea waybill data without templates

Lido uses a custom blend of AI vision models, OCR, and LLMs to extract data from documents in any format. You describe the fields you want in plain English, upload sample documents, and Lido returns structured data without requiring a template for every airline, forwarder, ocean carrier, or NVOCC.

For waybill processing, that means you can ask Lido to extract fields like:

  • AWB number, MAWB number, HAWB number, airline, origin airport, destination airport, flight routing, pieces, gross weight, chargeable weight, and charges
  • Sea waybill number, booking number, vessel, voyage, port of loading, port of discharge, container number, seal number, cargo description, package count, gross weight, measurement, and freight terms
  • Shipper, consignee, notify party, carrier, freight forwarder, customer reference, purchase order number, invoice number, and customs-related fields

Lido can export the extracted results to Excel, CSV, Google Sheets, JSON, or API. That matters because logistics teams rarely want another isolated OCR dashboard. They want the data to move into the system they already use: a freight management platform, TMS, customs brokerage workflow, shared spreadsheet, ERP, accounting system, or internal database.

The same approach also works across related documents. If a shipment packet includes a commercial invoice, packing list, bill of lading, delivery order, carrier invoice, or proof of delivery, Lido can extract the relevant fields from each document type instead of forcing your team to process every file manually. For customs-heavy workflows, see our comparison of the best customs document processing software.

What to look for in waybill OCR software

When you compare air waybill OCR software or sea waybill OCR tools, the feature list can sound similar. Most vendors will say they use AI, machine learning, OCR, or document AI. The practical test is whether the tool works on your actual shipping documents.

Use this checklist:

  • No per-carrier templates. You should not need a custom template for every airline, freight forwarder, ocean carrier, or NVOCC.
  • Support for MAWB and HAWB. Air freight workflows often need both master air waybills and house air waybills, sometimes matched together.
  • Support for ocean waybills and bills of lading. Sea waybills often sit next to bills of lading, delivery orders, packing lists, and commercial invoices in the same workflow.
  • Scanned and photographed document handling. The tool should work on forwarded PDFs, poor scans, skewed images, and documents with stamps or handwritten notes.
  • Custom fields. Your operation may need customer references, shipment IDs, PO numbers, internal lane codes, accessorial charges, or compliance fields that are not part of a generic schema.
  • Consistent output. Column names, date formats, weights, currency, and references should not change from file to file.
  • Validation and review workflow. High-risk fields should be checked before import when they are missing, unclear, or inconsistent.
  • Flexible exports. Look for Excel, CSV, Google Sheets, JSON, and API options so your team can connect the output to existing systems.
  • Fast iteration. You should be able to adjust extraction instructions and reprocess documents without rebuilding the workflow from scratch.

If you only have one perfectly consistent document format, a template tool might be enough. If you process many carriers, forwarders, and document types, a layout-agnostic approach like Lido is usually the better fit.

How to connect extracted waybill data to your workflow

Waybill OCR only creates value when the extracted data moves somewhere useful. Before choosing a tool, decide what the end state should be.

Common outputs include:

  • Excel or CSV for operations teams that import shipment data into legacy systems or reconcile documents manually.
  • Google Sheets for shared review queues, exception handling, and lightweight operations dashboards.
  • JSON for engineering teams that need structured document data paired with source files.
  • API for pushing waybill fields into a TMS, freight management system, customs brokerage platform, ERP, or internal workflow.
  • Human review queues for fields that should be checked before they affect clearance, release, billing, or customer communication.

A good implementation starts small. Pick a representative set of air waybills and sea waybills from different carriers. Define the fields you need. Decide which fields can flow straight through and which require review. Then connect the output to one downstream workflow before expanding to more document types.

That first workflow might be simple: upload AWBs and sea waybills, extract fields into a spreadsheet, review exceptions, and export CSV. Once the schema is stable, you can connect the same structured output to an API or internal system.

When Lido is a good fit

Lido is a good fit if your team processes shipping documents at enough volume that manual data entry is slowing down operations, but your document formats are too variable for rigid templates. That pattern shows up across logistics teams, including trucking companies processing handwritten BOLs, PODs, driver tickets, and carrier paperwork at scale.

It is especially useful when:

  • You process air waybills, MAWBs, HAWBs, sea waybills, bills of lading, invoices, packing lists, and other logistics documents together.
  • Your documents come from many carriers, forwarders, customers, and brokers.
  • You need to extract custom fields, not just a vendor’s default AWB or sea waybill schema.
  • Your source files include scanned PDFs, email attachments, phone photos, stamps, handwritten notes, or degraded images.
  • You need structured output in Excel, CSV, Google Sheets, JSON, or API format.
  • You want a workflow that can include human review for exceptions instead of full manual keying.

The documents do not need to become cleaner or more standardized for automation to work. The extraction layer needs to be flexible enough to handle the documents your team already receives.

Frequently asked questions

What is the best software for air waybill OCR and sea waybill OCR?

Lido is the best software for teams that need to extract structured data from air waybills, MAWBs, HAWBs, sea waybills, and related shipping documents without building templates for every carrier or forwarder. It handles variable layouts, scanned PDFs, and custom fields, then exports results to Excel, CSV, Google Sheets, JSON, or API.

What fields can air waybill OCR extract?

Air waybill OCR can extract AWB number, MAWB number, HAWB number, carrier, shipper, consignee, origin airport, destination airport, routing, flight number, pieces, gross weight, chargeable weight, commodity description, HS code, declared value, handling instructions, charges, insurance details, and shipment references.

What fields can sea waybill OCR extract?

Sea waybill OCR can extract sea waybill number, booking number, carrier reference, shipper, consignee, notify party, vessel, voyage, port of loading, port of discharge, place of receipt, place of delivery, container number, seal number, package count, cargo description, gross weight, measurement, charges, and special instructions.

Can OCR handle both MAWB and HAWB documents?

Yes. Lido can extract data from both master air waybills and house air waybills. It can pull fields from each document type and structure them separately, which is useful when a freight forwarder needs to match a MAWB for a consolidated shipment with one or more HAWBs for individual consignments.

Is a sea waybill the same as a bill of lading?

No. A sea waybill is generally a non-negotiable ocean freight document and is not a document of title. A bill of lading can be negotiable depending on the type and may function as a document of title. OCR software should classify these documents separately because the fields, release workflow, and compliance handling can differ.

Can waybill OCR export data to Excel, CSV, Google Sheets, or an API?

Yes. Lido can export extracted waybill data to Excel, CSV, Google Sheets, JSON, or API. Teams can use those outputs to update spreadsheets, feed a freight management system, support customs brokerage workflows, push data into a TMS, or connect with internal operations and accounting systems.

Is basic OCR enough for shipping document processing?

Basic OCR is usually not enough for shipping document processing because it only converts the page into text. Waybill automation needs document classification, field extraction, normalization, validation, and structured output. Lido handles those steps so teams can work with clean rows of data instead of unstructured OCR text.

Ready to grow your business with document automation, not headcount?

Join hundreds of teams growing faster by automating the busywork with Lido.