Logistics & Supply ChainParse → Extract

Receipt Extractor

Extracts itemized purchases, totals, payment details, and loyalty points from retail receipts.

Ship it with Extend

Live pipeline

a real document, processed end to end · view only
Source documentreceipt2.jpg

Step-by-step

A retail receipt is a point-of-sale document issued by a retail store that records itemized product purchases, transaction totals, payment method details, merchant information, and loyalty program activity for a completed transaction. This template takes in Retail Receipt and outputs markdown (.md) capturing the receipt's full text and layout, and JSON (.json) with structured transaction fields including retailer, items, totals, payment authorization, and loyalty program data per the extraction schema by using Extend's Parse, Extract primitives.

Input
Retail Receipt
Compatible document types (full list)
.pdf.docx.xlsx.png.jpg.jpeg.tiff.tif.svg.heic.heif.bmp.gif.webp.psd.xls.xltm.xltx.ods.doc.wpd.dotx.odt.pptx.ppt.ppm.csv.txt.html.xml.rtf.lis.md.eml.pcx
Step 1

Parse

Converts the document into clean, layout-aware markdown plus structured blocks with spatial metadata.

InputSource document — PDF, image, spreadsheet, presentation, or scan
Config
blockOptions.text.agentic.enabledtruechanged
chunkingStrategy.type"document"
engine"parse_performance"
OutputMarkdown chunked by page or section, plus typed blocks (text, table, figure) with bounding boxes

You can learn more about Parse configuration in Extend's Parse documentation.

Step 2

Extract

Pulls a defined set of fields from the document and returns them as structured JSON matching a schema.

InputOutput of the Parse step
Config
metadata.change_given.citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":198,"y":736},{"x":332,"y":738},{"x":331,"y":764},{"x":198,"y":761}],"pageWidth":873,"pageHeight…changed
metadata.change_given.logprobsConfidencenullchanged
metadata.change_given.ocrConfidence0.989changed
metadata.customer_name.citations[]changed
metadata.customer_name.logprobsConfidencenullchanged
metadata.discount_amount.citations[]changed
metadata.discount_amount.logprobsConfidencenullchanged
metadata.line_items.citations[]changed
metadata.line_items.logprobsConfidencenullchanged
metadata.line_items.ocrConfidence0.095changed
metadata.line_items[0].citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":198,"y":343},{"x":333,"y":347},{"x":333,"y":373},{"x":198,"y":369}],"pageWidth":873,"pageHeight…changed
metadata.line_items[0].description.logprobsConfidencenullchanged
metadata.line_items[0].description.ocrConfidence0.992changed
metadata.line_items[0].logprobsConfidencenullchanged
metadata.line_items[0].ocrConfidence0.095changed
metadata.line_items[0].quantity.logprobsConfidencenullchanged
metadata.line_items[0].quantity.ocrConfidence0.095changed
metadata.line_items[0].total_price.logprobsConfidencenullchanged
metadata.line_items[0].total_price.ocrConfidence0.987changed
metadata.line_items[0].unit_price.logprobsConfidencenullchanged
metadata.line_items[0].unit_price.ocrConfidence0.987changed
metadata.line_items[1].citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":198,"y":372},{"x":278,"y":373},{"x":278,"y":397},{"x":197,"y":396}],"pageWidth":873,"pageHeight…changed
metadata.line_items[1].description.logprobsConfidencenullchanged
metadata.line_items[1].description.ocrConfidence0.634changed
metadata.line_items[1].logprobsConfidencenullchanged
metadata.line_items[1].ocrConfidence0.095changed
metadata.line_items[1].quantity.logprobsConfidencenullchanged
metadata.line_items[1].quantity.ocrConfidence0.095changed
metadata.line_items[1].total_price.logprobsConfidencenullchanged
metadata.line_items[1].total_price.ocrConfidence0.992changed
metadata.line_items[1].unit_price.logprobsConfidencenullchanged
metadata.line_items[1].unit_price.ocrConfidence0.992changed
metadata.line_items[2].citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":198,"y":399},{"x":411,"y":405},{"x":410,"y":432},{"x":197,"y":425}],"pageWidth":873,"pageHeight…changed
metadata.line_items[2].description.logprobsConfidencenullchanged
metadata.line_items[2].description.ocrConfidence0.696changed
metadata.line_items[2].logprobsConfidencenullchanged
metadata.line_items[2].ocrConfidence0.541changed
metadata.line_items[2].quantity.logprobsConfidencenullchanged
metadata.line_items[2].quantity.ocrConfidence0.541changed
metadata.line_items[2].total_price.logprobsConfidencenullchanged
metadata.line_items[2].total_price.ocrConfidence0.992changed
metadata.line_items[2].unit_price.logprobsConfidencenullchanged
metadata.line_items[2].unit_price.ocrConfidence0.782changed
metadata.payment_method.citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":197,"y":511},{"x":399,"y":515},{"x":398,"y":540},{"x":197,"y":537}],"pageWidth":873,"pageHeight…changed
metadata.payment_method.logprobsConfidencenullchanged
metadata.payment_method.ocrConfidence0.992changed
metadata.receipt_number.citations[]changed
metadata.receipt_number.logprobsConfidencenullchanged
metadata.subtotal_amount.citations[]changed
metadata.subtotal_amount.logprobsConfidencenullchanged
metadata.tax_amount.citations[]changed
metadata.tax_amount.logprobsConfidencenullchanged
metadata.total_amount.citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":199,"y":483},{"x":265,"y":483},{"x":265,"y":509},{"x":198,"y":509}],"pageWidth":873,"pageHeight…changed
metadata.total_amount.logprobsConfidencenullchanged
metadata.total_amount.ocrConfidence0.992changed
metadata.transaction_date.citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":184,"y":1203},{"x":693,"y":1213},{"x":693,"y":1239},{"x":183,"y":1231}],"pageWidth":873,"pageHe…changed
metadata.transaction_date.logprobsConfidencenullchanged
metadata.transaction_date.ocrConfidence0.992changed
metadata.vendor_address.citations[]changed
metadata.vendor_address.logprobsConfidencenullchanged
metadata.vendor_email.citations[]changed
metadata.vendor_email.logprobsConfidencenullchanged
metadata.vendor_name.citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":371,"y":162},{"x":515,"y":165},{"x":514,"y":197},{"x":371,"y":193}],"pageWidth":873,"pageHeight…changed
metadata.vendor_name.logprobsConfidencenullchanged
metadata.vendor_name.ocrConfidence0.961changed
metadata.vendor_phone.citations[{"page":1,"fileId":"file_woDHCAZS1jB2zhLo3Ib8v","polygon":[{"x":350,"y":240},{"x":559,"y":251},{"x":558,"y":278},{"x":349,"y":267}],"pageWidth":873,"pageHeight…changed
metadata.vendor_phone.logprobsConfidencenullchanged
metadata.vendor_phone.ocrConfidence0.987changed
value.change_given.amount0changed
value.change_given.iso_4217_currency_code"GBP"changed
value.customer_namenullchanged
value.discount_amount.amountnullchanged
value.discount_amount.iso_4217_currency_codenullchanged
value.line_items[{"quantity":1,"unit_price":0.89,"description":"FRESH MILK","total_price":0.89},{"quantity":1,"unit_price":2.29,"description":"MUESLI","total_price":2.29},{"qua…changed
value.payment_method"Mastercard"changed
value.receipt_numbernullchanged
value.subtotal_amount.amountnullchanged
value.subtotal_amount.iso_4217_currency_codenullchanged
value.tax_amount.amountnullchanged
value.tax_amount.iso_4217_currency_codenullchanged
value.total_amount.amount5.08changed
value.total_amount.iso_4217_currency_code"GBP"changed
value.transaction_date"2012-04-23"changed
value.vendor_addressnullchanged
value.vendor_emailnullchanged
value.vendor_name"Tesco Metro"changed
value.vendor_phone"0845 6779218"changed
OutputJSON shaped to the extraction schema, with per-field confidence scores and citations grounding each value to its source location

You can learn more about Extract configuration in Extend's Extract documentation.

Example code

{
  "name": "Receipt Processing Pipeline",
  "steps": [
    {
      "name": "startTrigger1",
      "type": "TRIGGER",
      "next": [
        {
          "step": "parse1"
        }
      ]
    },
    {
      "name": "parse1",
      "type": "PARSE",
      "config": {
        "parseConfig": {
          "blockOptions": {
            "text": {
              "agentic": {
                "enabled": true
              }
            }
          },
          "chunkingStrategy": {
            "type": "document"
          }
        }
      },
      "next": [
        {
          "step": "extraction2"
        }
      ]
    },
    {
      "name": "extraction2",
      "type": "EXTRACT"
    }
  ]
}
# Receipt Processing — Extend AI Skill

## What this pipeline does

Extracts structured transaction data from retail receipts (supermarket, convenience store, chain retail) into itemized line items, totals, payment method, vendor details, and transaction timestamp. Uses agentic OCR to handle variable receipt layouts, thermal printer degradation, and crumpled/rotated receipts, then parses the markdown and extracts all fields into a strongly-typed JSON object with currency amounts, dates, and item arrays.

## When to use this

- **Expense report automation**: OCR receipts, extract line items and totals, feed to accounting systems
- **Receipt auditing**: Verify totals match itemization, detect discrepancies, flag missing tax/discount info
- **Loyalty program reconciliation**: Extract vendor name, transaction date, payment method to match against reward program records
- **Point-of-sale analytics**: Batch-process end-of-day receipts to track sales mix, pricing, tax compliance by location
- **Receipt image archival**: Parse receipt to markdown for full-text search while storing original image for dispute resolution

## Processor pipeline

### Step 1: Parse (agentic_ocr mode)
**Processor**: `parseRuns.createAndPoll()` with `blockOptions.text.agentic.enabled = true`

**Purpose**: Convert receipt image to markdown with bounding box citations. Agentic mode handles thermal print fade, rotation, creasing, and multi-column layouts.

**Config rationale**:
- `agentic: true` — receipts are high-variance (thermal degradation, wrinkled, rotated). Light mode fails on ~30% of real-world receipts.
- `chunkingStrategy: "document"` — receipts are short documents (1–3 pages); keep as single chunk to preserve line item order and table structure.
- No `outputType` override — defaults to markdown, which is ideal for agentic extraction.

**Output**: Markdown with text blocks, each tagged with page, bounding box polygon, and OCR confidence.

### Step 2: Extract (structured fields)
**Processor**: `extractRuns.createAndPoll()` with Zod schema

**Purpose**: Pull vendor, transaction, payment, and itemized data into typed JSON object. Zod schema enforces field types and provides field descriptions to guide the LLM's extraction accuracy.

**Config rationale**:
- **Line items as array of objects** — captures quantity, unit price, description, total per item; enables reconciliation and audit trails.
- **Currency fields** — use `extendCurrency()` helper to capture both numeric amount and ISO 4217 code (e.g., GBP, USD); critical for multi-currency workflows.
- **Nullable fields for optional data** — `vendor_email`, `customer_name`, `discount_amount` often absent; schema allows null to avoid hallucination.
- **ISO date for transaction_date** — standardizes date format across receipts (e.g., `"2012-04-23"`); simplifies downstream filtering and sorting.
- **Payment method as string** — captures card type (Mastercard, Visa, Amex) and payment type (cash, check, loyalty); enables reconciliation.

**Accuracy levers**:
- `description` fields are maximally specific: "Full legal name of the vendor on the receipt" vs generic "vendor name". LLM uses descriptions to disambiguate similar fields.
- Line item `description` includes "Product or service name as printed on receipt" to avoid capturing prices or SKUs.
- `total_amount` description explicitly says "including tax" to separate from pre-tax subtotal.

## TypeScript implementation



## CLI equivalent

```bash
# Step 1: Parse receipt to markdown (agentic OCR mode)
extend parse receipt.pdf \
  --config '{
    "blockOptions": {
      "text": { "agentic": { "enabled": true } }
    },
    "chunkingStrategy": { "type": "document" }
  }' \
  --output receipt_parsed.md

# Step 2: Extract structured fields (requires schema.json)
extend extract receipt.pdf \
  --schema schema.json \
  --output receipt_data.json
```

**schema.json** (for CLI):
```json
{
  "type": "object",
  "properties": {
    "vendor_name": {
      "type": ["string", "null"],
      "description": "Full legal name or trading name of the vendor/merchant on the receipt."
    },
    "vendor_address": {
      "type": ["string", "null"],
      "description": "Full street address of the retail location where the transaction occurred."
    },
    "vendor_phone": {
      "type": ["string", "null"],
      "description": "Primary phone number of the vendor/store."
    },
    "vendor_email": {
      "type": ["string", "null"],
      "description": "Email address of the vendor, if printed on receipt."
    },
    "receipt_number": {
      "type": ["string", "null"],
      "description": "Receipt or transaction number for identification."
    },
    "transaction_date": {
      "type": ["string", "null"],
      "description": "Date of the transaction in ISO 8601 format (yyyy-mm-dd)."
    },
    "payment_method": {
      "type": ["string", "null"],
      "description": "Method of payment (e.g., 'Mastercard', 'Visa', 'Cash')."
    },
    "subtotal_amount": {
      "type": "object",
      "properties": {
        "amount": { "type": ["number", "null"], "description": "Numeric amount" },
        "iso_4217_currency_code": { "type": ["string", "null"], "description": "ISO currency code (e.g., 'GBP')" }
      },
      "description": "Subtotal before tax and discounts."
    },
    "discount_amount": {
      "type": "object",
      "properties": {
        "amount": { "type": ["number", "null"] },
        "iso_4217_currency_code": { "type": ["string", "null"] }
      },
      "description": "Total discount or promotion amount applied."
    },
    "tax_amount": {
      "type": "object",
      "properties": {
        "amount": { "type": ["number", "null"] },
        "iso_4217_currency_code": { "type": ["string", "null"] }
      },
      "description": "Total sales tax or VAT charged."
    },
    "total_amount": {
      "type": "object",
      "properties": {
        "amount": { "type": ["number", "null"] },
        "iso_4217_currency_code": { "type": ["string", "null"] }
      },
      "description": "Final total amount due or paid, including all tax and after all discounts."
    },
    "change_given": {
      "type": "object",
      "properties": {
        "amount": { "type": ["number", "null"] },
        "iso_4217_currency_code": { "type": ["string", "null"] }
      },
      "description": "Change returned to customer if paid in cash."
    },
    "line_items": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "description": {
            "type": ["string", "null"],
            "description": "Product or service name/description as printed on receipt."
          },
          "quantity": {
            "type": ["number", "null"],
            "description": "Quantity purchased of this line item."
          },
          "unit_price": {
            "type": ["number", "null"],
            "description": "Price per unit."
          },
          "total_price": {
            "type": ["number", "null"],
            "description": "Total price for this line item."
          }
        }
      },
      "description": "Array of purchased items, each with description, quantity, unit price, and total."
    },
    "customer_name": {
      "type": ["string", "null"],
      "description": "Name of the customer, if printed on receipt."
    }
  }
}
```

Run the workflow directly:
```bash
# Execute a saved workflow (if you've stored this pipeline as workflow_rcpt123)
extend run
import { ExtendClient, extendDate, extendCurrency } from "extend-ai";
import { z } from "zod";
import fs from "fs";

const client = new ExtendClient({ token: process.env.EXTEND_API_KEY });

/**
 * Receipt Processing Pipeline
 *
 * 1. Parse receipt image to markdown using agentic OCR
 * 2. Extract structured fields (vendor, items, totals, payment) into typed JSON
 */
async function processReceipt(filePath: string) {
  // Convert local file to base64 data URL for SDK upload
  const fileBuffer = fs.readFileSync(filePath);
  const base64 = fileBuffer.toString("base64");
  const dataUrl = `data:application/octet-stream;base64,${base64}`;

  console.log(`[1/2] Parsing receipt from ${filePath} using agentic OCR...`);

  // Step 1: Parse receipt to markdown with bounding boxes
  const parseRun = await client.parseRuns.createAndPoll({
    file: { url: dataUrl },
    config: {
      blockOptions: {
        text: {
          agentic: {
            enabled: true, // Agentic mode handles thermal fade, rotation, creasing
          },
        },
      },
      chunkingStrategy: {
        type: "document", // Keep receipt as single chunk to preserve line item order
      },
    },
  });

  if (parseRun.status !== "PROCESSED") {
    throw new Error(`Parse failed with status: ${parseRun.status}`);
  }

  // Collect parsed markdown for inspection
  const parsedMarkdown = parseRun.output.chunks
    .map((chunk) => chunk.content)
    .join("\n\n");
  console.log(`[1/2] Parse complete. Output preview:\n${parsedMarkdown.slice(0, 500)}...\n`);

  console.log(`[2/2] Extracting receipt fields...`);

  // Step 2: Extract structured receipt data
  const extractRun = await client.extractRuns.createAndPoll({
    file: { url: dataUrl },
    config: {
      schema: z.object({
        vendor_name: z
          .string()
          .nullable()
          .describe(
            "Full legal name or trading name of the vendor/merchant on the receipt (e.g., 'Tesco Metro', 'Whole Foods'). Do not include store number or address."
          ),
        vendor_address: z
          .string()
          .nullable()
          .describe(
            "Full street address of the retail location where the transaction occurred, including city and postal code if present."
          ),
        vendor_phone: z
          .string()
          .nullable()
          .describe(
            "Primary phone number of the vendor/store, typically printed at top or bottom of receipt."
          ),
        vendor_email: z
          .string()
          .nullable()
          .describe("Email address of the vendor, if printed on receipt."),
        receipt_number: z
          .string()
          .nullable()
          .describe(
            "Receipt or transaction number (may be labeled as 'Receipt #', 'Trans ID', 'Reference', or similar). Used to identify this specific transaction."
          ),
        transaction_date: extendDate().describe(
          "Date of the transaction in ISO 8601 format (yyyy-mm-dd). Extract from date/time stamp on receipt."
        ),
        payment_method: z
          .string()
          .nullable()
          .describe(
            "Method of payment (e.g., 'Mastercard', 'Visa', 'American Express', 'Cash', 'Check', 'Apple Pay'). Include card brand if card is used."
          ),
        subtotal_amount: extendCurrency().describe(
          "Subtotal before tax and discounts. Labeled as 'Subtotal', 'Subtotal Ex Tax', or similar. Include both numeric amount and ISO 4217 currency code (e.g., 'GBP')."
        ),
        discount_amount: extendCurrency().describe(
          "Total discount or promotion amount applied, if any. Labeled as 'Discount', 'Promo', 'Loyalty Discount', or similar. Null if no discount."
        ),
        tax_amount: extendCurrency().describe(
          "Total sales tax or VAT charged. Labeled as 'Tax', 'VAT', 'Sales Tax', or similar. Null if tax-exempt or not shown separately."
        ),
        total_amount: extendCurrency().describe(
          "Final total amount due or paid, including all tax and after all discounts. This is the 'TOTAL' line. Include both numeric amount and ISO 4217 currency code."
        ),
        change_given: extendCurrency().describe(
          "Change returned to customer if paid in cash. Labeled as 'Change', 'Change Due', or 'Change Given'. Null if card payment or exact cash."
        ),
        line_items: z
          .array(
            z.object({
              description: z
                .string()
                .nullable()
                .describe(
                  "Product or service name/description as printed on receipt (e.g., 'FRESH MILK', 'DARK CHOCOLATE'). Do not include price, quantity, or SKU."
                ),
              quantity: z
                .number()
                .nullable()
                .describe(
                  "Quantity purchased of this line item. If quantity is not explicit, default to 1. Handle '2 @' notation (2 units at unit price)."
                ),
              unit_price: z
                .number()
                .nullable()
                .describe(
                  "Price per unit (not total price). If only total price is shown and quantity is known, divide to get unit price. Otherwise, unit price equals total price."
                ),
              total_price: z
                .number()
                .nullable()
                .describe(
                  "Total price for this line item (quantity × unit_price, before tax if tax is itemized). This is the amount shown on the right side of the line."
                ),
            })
          )
          .describe(
            "Array of purchased items. Each item includes description, quantity, unit price, and line total. Preserve order as shown on receipt."
          ),
        customer_name: z
          .string()
          .nullable()
          .describe(
            "Name of the customer, if printed on receipt (e.g., loyalty member name, credit card holder name). Usually not present on receipts; null if absent."
          ),
      }),
    },
  });

  if (extractRun.status !== "PROCESSED") {
    throw new Error(`Extract failed with status: ${extractRun.status}`);
  }

  const receipt = extractRun.output.value;

  console.log(`[2/2] Extraction complete.\n`);
  console.log("=== RECEIPT DATA ===");
  console.log(`Vendor: ${receipt.vendor_name}`);
  console.log(`Date: ${receipt.transaction_date}`);
  console.log(`Total: ${receipt.total_amount?.amount} ${receipt.total_amount?.iso_4217_currency_code}`);
  console.log(`Payment: ${receipt.payment_method}`);
  console.log(`\nLine Items:`);
  receipt.line_items.forEach((item, i) => {
    console.log(
      `  ${i + 1}. ${item.description} x${item.quantity} @ ${item.unit_price} = ${item.total_price}`
    );
  });
  console.log(`\nSubtotal: ${receipt.subtotal_amount?.amount}`);
  console.log(`Tax: ${receipt.tax_amount?.amount}`);
  console.log(`Discount: ${receipt.discount_amount?.amount}`);
  console.log(`Change Given: ${receipt.change_given?.amount}`);

  return receipt;
}

// Export for test harness
export { processReceipt };

// Allow direct execution for debugging
if (require.main === module) {
  const testFile = process.argv[2] || "./receipt.pdf";
  processReceipt(testFile)
    .then(() => console.log("\n✓ Processing complete"))
    .catch((err) => {
      console.error("\n✗ Processing failed:", err.message);
      process.exit(1);
    });
}
import os
import base64
from extend_ai import Extend

client = Extend(token=os.environ["EXTEND_API_KEY"])


def process_receipt(file_path: str):
    """
    Receipt Processing Pipeline

    1. Parse receipt image to markdown using agentic OCR
    2. Extract structured fields (vendor, items, totals, payment) into typed JSON
    """
    # Convert local file to base64 data URL for SDK upload
    with open(file_path, "rb") as f:
        file_buffer = f.read()
    base64_str = base64.b64encode(file_buffer).decode("utf-8")
    data_url = f"data:application/octet-stream;base64,{base64_str}"

    print(f"[1/2] Parsing receipt from {file_path} using agentic OCR...")

    # Step 1: Parse receipt to markdown with bounding boxes
    parse_run = client.parse_runs.create_and_poll(
        file={"url": data_url},
        config={
            "blockOptions": {
                "text": {
                    "agentic": {
                        "enabled": True,  # Agentic mode handles thermal fade, rotation, creasing
                    },
                },
            },
            "chunkingStrategy": {
                "type": "document",  # Keep receipt as single chunk to preserve line item order
            },
        },
    )

    if parse_run.status != "PROCESSED":
        raise Exception(f"Parse failed with status: {parse_run.status}")

    # Collect parsed markdown for inspection
    parsed_markdown = "\n\n".join(chunk.content for chunk in parse_run.output.chunks)
    print(f"[1/2] Parse complete. Output preview:\n{parsed_markdown[:500]}...\n")

    print("[2/2] Extracting receipt fields...")

    # Step 2: Extract structured receipt data
    extract_run = client.extract_runs.create_and_poll(
        file={"url": data_url},
        config={
            "schema": {
                "type": "object",
                "properties": {
                    "retailer": {
                        "type": ["string", "null"],
                        "description": "Name of the retail store",
                    },
                    "transaction_date": {
                        "type": ["string", "null"],
                        "description": "Date of transaction in DD/MM/YY format",
                    },
                    "transaction_time": {
                        "type": ["string", "null"],
                        "description": "Time of transaction in HH:MM format",
                    },
                    "items": {
                        "type": "array",
                        "description": "List of items purchased",
                        "items": {
                            "type": "object",
                            "properties": {
                                "name": {
                                    "type": ["string", "null"],
                                    "description": "Product name",
                                },
                                "price": {
                                    "type": ["number", "null"],
                                    "description": "Item price in GBP",
                                },
                                "quantity": {
                                    "type": ["number", "null"],
                                    "description": "Quantity purchased",
                                },
                            },
                        },
                    },
                    "subtotal": {
                        "type": ["number", "null"],
                        "description": "Subtotal amount in GBP",
                    },
                    "total": {
                        "type": ["number", "null"],
                        "description": "Total amount paid in GBP",
                    },
                    "payment_method": {
                        "type": ["string", "null"],
                        "description": "Payment method used",
                    },
                    "card_last_four": {
                        "type": ["string", "null"],
                        "description": "Last four digits of card number",
                    },
                    "auth_code": {
                        "type": ["string", "null"],
                        "description": "Authorization code for transaction",
                    },
                    "clubcard_number": {
                        "type": ["string", "null"],
                        "description": "Loyalty card number (masked)",
                    },
                    "points_earned": {
                        "type": ["number", "null"],
                        "description": "Loyalty points earned this visit",
                    },
                    "merchant_id": {
                        "type": ["string", "null"],
                        "description": "Merchant identification number",
                    },
                },
            },
        },
    )

    if extract_run.status != "PROCESSED":
        raise Exception(f"Extract failed with status: {extract_run.status}")

    receipt = extract_run.output.value

    print("[2/2] Extraction complete.\n")
    print("=== RECEIPT DATA ===")
    print(f"Retailer: {receipt.get('retailer')}")
    print(f"Date: {receipt.get('transaction_date')}")
    print(f"Time: {receipt.get('transaction_time')}")
    print(f"Total: £{receipt.get('total')}")
    print(f"Payment: {receipt.get('payment_method')}")
    print("\nLine Items:")
    for i, item in enumerate(receipt.get("items", []), 1):
        print(
            f"  {i}. {item.get('name')} x{item.get('quantity')} @ £{item.get('price')}"
        )
    print(f"\nSubtotal: £{receipt.get('subtotal')}")
    print(f"Card Last Four: {receipt.get('card_last_four')}")
    print(f"Auth Code: {receipt.get('auth_code')}")
    print(f"Clubcard Number: {receipt.get('clubcard_number')}")
    print(f"Points Earned: {receipt.get('points_earned')}")
    print(f"Merchant ID: {receipt.get('merchant_id')}")

    return receipt


if __name__ == "__main__":
    import sys

    test_file = sys.argv[1] if len(sys.argv) > 1 else "./receipt.pdf"
    try:
        process_receipt(test_file)
        print("\n✓ Processing complete")
    except Exception as err:
        print(f"\n✗ Processing failed: {err}")
        sys.exit(1)
// This code uses Extend's REST API directly because Extend has no official Java SDK yet.
// It calls https://api.extend.ai endpoints with java.net.http.HttpClient (no external dependencies).

import java.io.IOException;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Base64;
import java.util.Scanner;

public class ReceiptProcessor {
  private static final String API_BASE = "https://api.extend.ai";
  private static final String API_KEY = System.getenv("EXTEND_API_KEY");
  private static final HttpClient httpClient = HttpClient.newHttpClient();

  /**
   * Receipt Processing Pipeline
   *
   * 1. Parse receipt image to markdown using agentic OCR
   * 2. Extract structured fields (vendor, items, totals, payment) into typed JSON
   */
  public static void processReceipt(String filePath) throws IOException, InterruptedException {
    // Convert local file to base64 data URL for API upload
    byte[] fileBytes = Files.readAllBytes(Paths.get(filePath));
    String base64 = Base64.getEncoder().encodeToString(fileBytes);
    String dataUrl = "data:application/octet-stream;base64," + base64;

    System.out.println("[1/2] Parsing receipt from " + filePath + " using agentic OCR...");

    // Step 1: Parse receipt to markdown with bounding boxes
    String parsePayload = "{"
        + "\"file\":{\"url\":\"" + escapeJson(dataUrl) + "\"},"
        + "\"config\":{"
        + "\"blockOptions\":{\"text\":{\"agentic\":{\"enabled\":true}}},"
        + "\"chunkingStrategy\":{\"type\":\"document\"}"
        + "}"
        + "}";

    String parseRunId = createAndPollParseRun(parsePayload);

    // Fetch parse run result
    String parseResult = fetchParseRunResult(parseRunId);
    String parsedMarkdown = extractMarkdownFromParseResult(parseResult);
    System.out.println("[1/2] Parse complete. Output preview:\n"
        + parsedMarkdown.substring(0, Math.min(500, parsedMarkdown.length())) + "...\n");

    System.out.println("[2/2] Extracting receipt fields...");

    // Step 2: Extract structured receipt data
    String schema = "{"
        + "\"type\":\"object\","
        + "\"properties\":{"
        + "\"retailer\":{\"type\":[\"string\",\"null\"],\"description\":\"Name of the retail store\"},"
        + "\"transaction_date\":{\"type\":[\"string\",\"null\"],\"description\":\"Date of transaction in DD/MM/YY format\"},"
        + "\"transaction_time\":{\"type\":[\"string\",\"null\"],\"description\":\"Time of transaction in HH:MM format\"},"
        + "\"items\":{\"type\":\"array\",\"description\":\"List of items purchased\",\"items\":{\"type\":\"object\",\"properties\":{\"name\":{\"type\":[\"string\",\"null\"],\"description\":\"Product name\"},\"price\":{\"type\":[\"number\",\"null\"],\"description\":\"Item price in GBP\"},\"quantity\":{\"type\":[\"number\",\"null\"],\"description\":\"Quantity purchased\"}}}},"
        + "\"subtotal\":{\"type\":[\"number\",\"null\"],\"description\":\"Subtotal amount in GBP\"},"
        + "\"total\":{\"type\":[\"number\",\"null\"],\"description\":\"Total amount paid in GBP\"},"
        + "\"payment_method\":{\"type\":[\"string\",\"null\"],\"description\":\"Payment method used\"},"
        + "\"card_last_four\":{\"type\":[\"string\",\"null\"],\"description\":\"Last four digits of card number\"},"
        + "\"auth_code\":{\"type\":[\"string\",\"null\"],\"description\":\"Authorization code for transaction\"},"
        + "\"clubcard_number\":{\"type\":[\"string\",\"null\"],\"description\":\"Loyalty card number (masked)\"},"
        + "\"points_earned\":{\"type\":[\"number\",\"null\"],\"description\":\"Loyalty points earned this visit\"},"
        + "\"merchant_id\":{\"type\":[\"string\",\"null\"],\"description\":\"Merchant identification number\"}"
        + "}"
        + "}";

    String extractPayload = "{"
        + "\"file\":{\"url\":\"" + escapeJson(dataUrl) + "\"},"
        + "\"config\":{\"schema\":" + schema + "}"
        + "}";

    String extractRunId = createAndPollExtractRun(extractPayload);

    // Fetch extract run result
    String extractResult = fetchExtractRunResult(extractRunId);
    printReceiptData(extractResult);

    System.out.println("\n✓ Processing complete");
  }

  private static String createAndPollParseRun(String payload)
      throws IOException, InterruptedException {
    HttpRequest request = HttpRequest.newBuilder()
        .uri(URI.create(API_BASE + "/v1/parseRuns"))
        .header("Authorization", "Bearer " + API_KEY)
        .header("Content-Type", "application/json")
        .POST(HttpRequest.BodyPublishers.ofString(payload))
        .build();

    HttpResponse<String> response = httpClient.send(request, HttpResponse.BodyHandlers.ofString());
    String runId = extractFieldFromJson(response.body(), "id");

    // Poll until PROCESSED
    while (true) {
      Thread.sleep(2000);
      HttpRequest statusRequest = HttpRequest.newBuilder()
          .uri(URI.create(API_BASE + "/v1/parseRuns/" + runId))
          .header("Authorization", "Bearer " + API_KEY)
          .GET()
          .build();

      HttpResponse<String> statusResponse = httpClient.send(statusRequest,
          HttpResponse.BodyHandlers.ofString());
      String status = extractFieldFromJson(statusResponse.body(), "status");

      if ("PROCESSED".equals(status)) {
        break;
      } else if ("FAILED".equals(status) || "ERROR".equals(status)) {
        throw new RuntimeException("Parse run failed with status: " + status);
      }
    }

    return runId;
  }

  private static String createAndPollExtractRun(String payload)
      throws IOException, InterruptedException {
    HttpRequest request = HttpRequest.newBuilder()
        .uri(URI.create(API_BASE + "/v1/extractRuns"))
        .header("Authorization", "Bearer " + API_KEY)
        .header("Content-Type", "application/json")
        .POST(HttpRequest.BodyPublishers.ofString(payload))
        .build();

    HttpResponse<String> response = httpClient.send(request, HttpResponse.BodyHandlers.ofString());
    String runId = extractFieldFromJson(response.body(), "id");

    // Poll until PROCESSED
    while (true) {
      Thread.sleep(2000);
      HttpRequest statusRequest = HttpRequest.newBuilder()
          .uri(URI.create(API_BASE + "/v1/extractRuns/" + runId))
          .header("Authorization", "Bearer " + API_KEY)
          .GET()
          .build();

      HttpResponse<String> statusResponse = httpClient.send(statusRequest,
          HttpResponse.BodyHandlers.ofString());
      String status = extractFieldFromJson(statusResponse.body(), "status");

      if ("PROCESSED".equals(status)) {
        break;
      } else if ("FAILED".equals(status) || "ERROR".equals(status)) {
        throw new RuntimeException("Extract run failed with status: " + status);
      }
    }

    return runId;
  }

  private static String fetchParseRunResult(String runId) throws IOException, InterruptedException {
    HttpRequest request = HttpRequest.newBuilder()
        .uri(URI.create(API_BASE + "/v1/parseRuns/" + runId))
        .header("Authorization", "Bearer " + API_KEY)
        .GET()
        .build();

    HttpResponse<String> response = httpClient.send(request, HttpResponse.BodyHandlers.ofString());
    return response.body();
  }

  private static String fetchExtractRunResult(String runId)
      throws IOException, InterruptedException {
    HttpRequest request = HttpRequest.newBuilder()
        .uri(URI.create(API_BASE + "/v1/extractRuns/" + runId))
        .header("Authorization", "Bearer " + API_KEY)
        .GET()
        .build();

    HttpResponse<String> response = httpClient.send(request, HttpResponse.BodyHandlers.ofString());
    return response.body();
  }

  private static String extractMarkdownFromParseResult(String json) {
    // Simple extraction of chunks content from parse result
    int chunksIdx = json.indexOf("\"chunks\"");
    if (chunksIdx == -1)
      return "";
    int contentIdx = json.indexOf("\"content\"", chunksIdx);
    if (contentIdx == -1)
      return "";
    int startQuote = json.indexOf("\"", contentIdx + 9);
    int endQuote = json.indexOf("\"", startQuote + 1);
    if (startQuote == -1 || endQuote == -1)
      return "";
    return json.substring(startQuote + 1, endQuote).replace("\\n", "\n");
  }

  private static void printReceiptData(String json) {
    System.out.println("[2/2] Extraction complete.\n");
    System.out.println("=== RECEIPT DATA ===");

    String retailer = extractFieldFromJson(json, "retailer");
    String transactionDate = extractFieldFromJson(json, "transaction_date");
    String total = extractFieldFromJson(json, "total");
    String paymentMethod = extractFieldFromJson(json, "payment_method");
    String subtotal = extractFieldFromJson(json, "subtotal");
    String authCode = extractFieldFromJson(json, "auth_code");
    String clubcard = extractFieldFromJson(json, "clubcard_number");

    System.out.println("Retailer: " + (retailer != null ? retailer : "N/A"));
    System.out.println("Date: " + (transactionDate != null ? transactionDate : "N/A"));
    System.out.println("Total: " + (total != null ? total : "N/A"));
    System.out.println("Payment: " + (paymentMethod != null ? paymentMethod : "N/A"));
    System.out.println("\nSubtotal: " + (subtotal != null ? subtotal : "N/A"));
    System.out.println("Auth Code: " + (authCode != null ? authCode : "N/A"));
    System.out.println("Clubcard: " + (clubcard != null ? clubcard : "N/A"));
  }

  private static String extractFieldFromJson(String json, String fieldName) {
    String searchKey = "\"" + fieldName + "\":";
    int idx = json.indexOf(searchKey);
    if (idx == -1)
      return null;

    int startIdx = idx + searchKey.length();
    while (startIdx < json.length() && Character.isWhitespace(json.charAt(startIdx))) {
      startIdx++;
    }

    if (startIdx >= json.length())
      return null;

    if (json.charAt(startIdx) == '"') {
      int endIdx = json.indexOf("\"", startIdx + 1);
      if (endIdx == -1)
        return null;
      return json.substring(startIdx + 1, endIdx);
    } else if (json.charAt(startIdx) == 'n') {
      return null;
    } else {
      int endIdx = startIdx;
      while (endIdx < json.length() && json.charAt(endIdx) != ',' && json.charAt(endIdx) != '}') {
        endIdx++;
      }
      return json.substring(startIdx, endIdx).trim();
    }
  }

  private static String escapeJson(String str) {
    return str.replace("\\", "\\\\").replace("\"", "\\\"").replace("\n", "\\n")
        .replace("\r", "\\r").replace("\t", "\\t");
  }

  public static void main(String[] args) {
    try {
      String testFile = args.length > 0 ? args[0] : "./receipt.pdf";
      processReceipt(testFile);
    } catch (Exception e) {
      System.err.println("\n✗ Processing failed: " + e.getMessage());
      e.printStackTrace();
      System.exit(1);
    }
  }
}
// This code uses the Extend REST API directly because Extend has no official Go SDK yet.
// It calls https://api.extend.ai endpoints with standard net/http and encoding/json.

package main

import (
	"bytes"
	"encoding/base64"
	"encoding/json"
	"fmt"
	"io"
	"net/http"
	"os"
	"time"
)

const extendAPIBase = "https://api.extend.ai"

type ExtendClient struct {
	token string
}

func NewExtendClient(token string) *ExtendClient {
	return &ExtendClient{token: token}
}

// ParseRunConfig defines the configuration for parse operations
type ParseRunConfig struct {
	BlockOptions struct {
		Text struct {
			Agentic struct {
				Enabled bool `json:"enabled"`
			} `json:"agentic"`
		} `json:"text"`
	} `json:"blockOptions"`
	ChunkingStrategy struct {
		Type string `json:"type"`
	} `json:"chunkingStrategy"`
}

// ParseRunRequest is the request body for creating a parse run
type ParseRunRequest struct {
	File struct {
		URL string `json:"url"`
	} `json:"file"`
	Config ParseRunConfig `json:"config"`
}

// ParseRunResponse is the response from a parse run
type ParseRunResponse struct {
	ID     string `json:"id"`
	Status string `json:"status"`
	Output struct {
		Chunks []struct {
			Content string `json:"content"`
		} `json:"chunks"`
	} `json:"output"`
}

// ExtractRunConfig defines the schema for extraction
type ExtractRunConfig struct {
	Schema map[string]interface{} `json:"schema"`
}

// ExtractRunRequest is the request body for creating an extract run
type ExtractRunRequest struct {
	File struct {
		URL string `json:"url"`
	} `json:"file"`
	Config ExtractRunConfig `json:"config"`
}

// ExtractRunResponse is the response from an extract run
type ExtractRunResponse struct {
	ID     string `json:"id"`
	Status string `json:"status"`
	Output struct {
		Value map[string]interface{} `json:"value"`
	} `json:"output"`
}

// LineItem represents a purchased item
type LineItem struct {
	Name     *string  `json:"name"`
	Price    *float64 `json:"price"`
	Quantity *float64 `json:"quantity"`
}

// Receipt represents the extracted receipt data
type Receipt struct {
	Retailer      *string     `json:"retailer"`
	TransactionDate *string   `json:"transaction_date"`
	TransactionTime *string   `json:"transaction_time"`
	Items          []LineItem `json:"items"`
	Subtotal       *float64   `json:"subtotal"`
	Total          *float64   `json:"total"`
	PaymentMethod  *string    `json:"payment_method"`
	CardLastFour   *string    `json:"card_last_four"`
	AuthCode       *string    `json:"auth_code"`
	ClubcardNumber *string    `json:"clubcard_number"`
	PointsEarned   *float64   `json:"points_earned"`
	MerchantID     *string    `json:"merchant_id"`
}

// createAndPollParseRun creates a parse run and polls until completion
func (c *ExtendClient) createAndPollParseRun(req ParseRunRequest) (*ParseRunResponse, error) {
	body, err := json.Marshal(req)
	if err != nil {
		return nil, err
	}

	httpReq, err := http.NewRequest("POST", extendAPIBase+"/v1/parseRuns", bytes.NewReader(body))
	if err != nil {
		return nil, err
	}

	httpReq.Header.Set("Authorization", "Bearer "+c.token)
	httpReq.Header.Set("Content-Type", "application/json")

	client := &http.Client{}
	resp, err := client.Do(httpReq)
	if err != nil {
		return nil, err
	}
	defer resp.Body.Close()

	respBody, err := io.ReadAll(resp.Body)
	if err != nil {
		return nil, err
	}

	var parseResp ParseRunResponse
	if err := json.Unmarshal(respBody, &parseResp); err != nil {
		return nil, err
	}

	// Poll until completion
	for parseResp.Status != "PROCESSED" && parseResp.Status != "FAILED" {
		time.Sleep(2 * time.Second)

		pollReq, err := http.NewRequest("GET", extendAPIBase+"/v1/parseRuns/"+parseResp.ID, nil)
		if err != nil {
			return nil, err
		}
		pollReq.Header.Set("Authorization", "Bearer "+c.token)

		pollResp, err := client.Do(pollReq)
		if err != nil {
			return nil, err
		}
		defer pollResp.Body.Close()

		pollBody, err := io.ReadAll(pollResp.Body)
		if err != nil {
			return nil, err
		}

		if err := json.Unmarshal(pollBody, &parseResp); err != nil {
			return nil, err
		}
	}

	return &parseResp, nil
}

// createAndPollExtractRun creates an extract run and polls until completion
func (c *ExtendClient) createAndPollExtractRun(req ExtractRunRequest) (*ExtractRunResponse, error) {
	body, err := json.Marshal(req)
	if err != nil {
		return nil, err
	}

	httpReq, err := http.NewRequest("POST", extendAPIBase+"/v1/extractRuns", bytes.NewReader(body))
	if err != nil {
		return nil, err
	}

	httpReq.Header.Set("Authorization", "Bearer "+c.token)
	httpReq.Header.Set("Content-Type", "application/json")

	client := &http.Client{}
	resp, err := client.Do(httpReq)
	if err != nil {
		return nil, err
	}
	defer resp.Body.Close()

	respBody, err := io.ReadAll(resp.Body)
	if err != nil {
		return nil, err
	}

	var extractResp ExtractRunResponse
	if err := json.Unmarshal(respBody, &extractResp); err != nil {
		return nil, err
	}

	// Poll until completion
	for extractResp.Status != "PROCESSED" && extractResp.Status != "FAILED" {
		time.Sleep(2 * time.Second)

		pollReq, err := http.NewRequest("GET", extendAPIBase+"/v1/extractRuns/"+extractResp.ID, nil)
		if err != nil {
			return nil, err
		}
		pollReq.Header.Set("Authorization", "Bearer "+c.token)

		pollResp, err := client.Do(pollReq)
		if err != nil {
			return nil, err
		}
		defer pollResp.Body.Close()

		pollBody, err := io.ReadAll(pollResp.Body)
		if err != nil {
			return nil, err
		}

		if err := json.Unmarshal(pollBody, &extractResp); err != nil {
			return nil, err
		}
	}

	return &extractResp, nil
}

// ProcessReceipt processes a receipt image through parse and extract pipeline
func ProcessReceipt(filePath string) (*Receipt, error) {
	// Read file and convert to base64 data URL
	fileBuffer, err := os.ReadFile(filePath)
	if err != nil {
		return nil, err
	}

	base64Str := base64.StdEncoding.EncodeToString(fileBuffer)
	dataURL := "data:application/octet-stream;base64," + base64Str

	client := NewExtendClient(os.Getenv("EXTEND_API_KEY"))

	fmt.Printf("[1/2] Parsing receipt from %s using agentic OCR...\n", filePath)

	// Step 1: Parse receipt to markdown
	parseReq := ParseRunRequest{}
	parseReq.File.URL = dataURL
	parseReq.Config.BlockOptions.Text.Agentic.Enabled = true
	parseReq.Config.ChunkingStrategy.Type = "document"

	parseRun, err := client.createAndPollParseRun(parseReq)
	if err != nil {
		return nil, err
	}

	if parseRun.Status != "PROCESSED" {
		return nil, fmt.Errorf("parse failed with status: %s", parseRun.Status)
	}

	// Collect parsed markdown
	var parsedMarkdown string
	for i, chunk := range parseRun.Output.Chunks {
		if i > 0 {
			parsedMarkdown += "\n\n"
		}
		parsedMarkdown += chunk.Content
	}

	if len(parsedMarkdown) > 500 {
		fmt.Printf("[1/2] Parse complete. Output preview:\n%s...\n\n", parsedMarkdown[:500])
	} else {
		fmt.Printf("[1/2] Parse complete. Output preview:\n%s\n\n", parsedMarkdown)
	}

	fmt.Println("[2/2] Extracting receipt fields...")

	// Step 2: Extract structured receipt data
	schema := map[string]interface{}{
		"type": "object",
		"properties": map[string]interface{}{
			"retailer": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Name of the retail store",
			},
			"transaction_date": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Date of transaction in DD/MM/YY format",
			},
			"transaction_time": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Time of transaction in HH:MM format",
			},
			"items": map[string]interface{}{
				"type":        "array",
				"description": "List of items purchased",
				"items": map[string]interface{}{
					"type": "object",
					"properties": map[string]interface{}{
						"name": map[string]interface{}{
							"type":        []string{"string", "null"},
							"description": "Product name",
						},
						"price": map[string]interface{}{
							"type":        []string{"number", "null"},
							"description": "Item price in GBP",
						},
						"quantity": map[string]interface{}{
							"type":        []string{"number", "null"},
							"description": "Quantity purchased",
						},
					},
				},
			},
			"subtotal": map[string]interface{}{
				"type":        []string{"number", "null"},
				"description": "Subtotal amount in GBP",
			},
			"total": map[string]interface{}{
				"type":        []string{"number", "null"},
				"description": "Total amount paid in GBP",
			},
			"payment_method": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Payment method used",
			},
			"card_last_four": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Last four digits of card number",
			},
			"auth_code": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Authorization code for transaction",
			},
			"clubcard_number": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Loyalty card number (masked)",
			},
			"points_earned": map[string]interface{}{
				"type":        []string{"number", "null"},
				"description": "Loyalty points earned this visit",
			},
			"merchant_id": map[string]interface{}{
				"type":        []string{"string", "null"},
				"description": "Merchant identification number",
			},
		},
	}

	extractReq := ExtractRunRequest{}
	extractReq.File.URL = dataURL
	extractReq.Config.Schema = schema

	extractRun, err := client.createAndPollExtractRun(extractReq)
	if err != nil {
		return nil, err
	}

	if extractRun.Status != "PROCESSED" {
		return nil, fmt.Errorf("extract failed with status: %s", extractRun.Status)
	}

	// Parse the extracted value into Receipt struct
	receiptJSON, err := json.Marshal(extractRun.Output.Value)
	if err != nil {
		return nil, err
	}

	var receipt Receipt
	if err := json.Unmarshal(receiptJSON, &receipt); err != nil {
		return nil, err
	}

	fmt.Println("[2/2] Extraction complete.\n")
	fmt.Println("=== RECEIPT DATA ===")
	if receipt.Retailer != nil {
		fmt.Printf("Retailer: %s\n", *receipt.Retailer)
	}
	if receipt.TransactionDate != nil {
		fmt.Printf("Date: %s\n", *receipt.TransactionDate)
	}
	if receipt.Total != nil {
		fmt.Printf("Total: £%.2f\n", *receipt.Total)
	}
	if receipt.PaymentMethod != nil {
		fmt.Printf("Payment: %s\n", *receipt.PaymentMethod)
	}

	fmt.Println("\nLine Items:")
	for i, item := range receipt.Items {
		qty := "1"
		if item.Quantity != nil {
			qty = fmt.Sprintf("%.0f", *item.Quantity)
		}
		price := "0.00"
		if item.Price != nil {
			price = fmt.Sprintf("%.2f", *item.Price)
		}
		name := ""
		if item.Name != nil {
			name = *item.Name
		}
		fmt.Printf("  %d. %s x%s @ £%s\n", i+1, name, qty, price)
	}

	if receipt.Subtotal != nil {
		fmt.Printf("\nSubtotal: £%.2f\n", *receipt.Subtotal)
	}
	if receipt.CardLastFour != nil {
		fmt.Printf("Card Last Four: %s\n", *receipt.CardLastFour)
	}
	if receipt.PointsEarned != nil {
		fmt.Printf("Points Earned: %.0f\n", *receipt.PointsEarned)
	}

	return &receipt, nil
}

func main() {
	testFile := "./receipt.pdf"
	if len(os.Args) > 1 {
		testFile = os.Args[1]
	}

	_, err := ProcessReceipt(testFile)
	if err != nil {
		fmt.Printf("\n✗ Processing failed: %v\n", err)
		os.Exit(1)
	}

	fmt.Println("\n✓ Processing complete")
}
// Deploy the "Receipt" pipeline to YOUR Extend account.
//
// The workflow below is fully self-contained — every EXTRACT/CLASSIFY/SPLIT
// step carries its extractor/classifier/splitter config INLINE, so this is a
// single API call. No processors to create or wire up beforehand.
// Idempotent: the created workflow id is cached in .extend/receipt-extractor.json,
// so re-running updates the existing workflow instead of duplicating it.
//
// Usage:
//   export EXTEND_API_KEY=sk_...   (from https://dashboard.extend.ai → API Keys)
//   npx tsx provision.ts
//
// Generated by doc1 (template: receipt-extractor).

import fs from "node:fs";
import path from "node:path";

const API = "https://api.extend.ai";
const VERSION = "2026-02-09";
const API_KEY = process.env.EXTEND_API_KEY;
if (!API_KEY) { console.error("Set EXTEND_API_KEY first."); process.exit(1); }

const STATE_DIR = path.join(process.cwd(), ".extend");
const STATE_FILE = path.join(STATE_DIR, "receipt-extractor.json");

type State = { workflowId?: string };
const state: State = fs.existsSync(STATE_FILE)
  ? JSON.parse(fs.readFileSync(STATE_FILE, "utf8"))
  : {};
function saveState() {
  fs.mkdirSync(STATE_DIR, { recursive: true });
  fs.writeFileSync(STATE_FILE, JSON.stringify(state, null, 2));
}

async function api(method: string, pathName: string, body?: unknown) {
  const res = await fetch(API + pathName, {
    method,
    headers: {
      Authorization: `Bearer ${API_KEY}`,
      "x-extend-api-version": VERSION,
      ...(body ? { "Content-Type": "application/json" } : {}),
    },
    body: body ? JSON.stringify(body) : undefined,
  });
  const data = await res.json().catch(() => ({}));
  if (!res.ok) throw new Error(`${method} ${pathName} failed (${res.status}): ${JSON.stringify(data).slice(0, 300)}`);
  return data;
}

// ── Workflow definition — extractor/classifier/splitter configs inline ──────
const WORKFLOW = {
  "name": "Receipt Processing Pipeline",
  "steps": [
    {
      "name": "startTrigger1",
      "type": "TRIGGER",
      "next": [
        {
          "step": "parse1"
        }
      ]
    },
    {
      "name": "parse1",
      "type": "PARSE",
      "config": {
        "parseConfig": {
          "blockOptions": {
            "text": {
              "agentic": {
                "enabled": true
              }
            }
          },
          "chunkingStrategy": {
            "type": "document"
          }
        }
      },
      "next": [
        {
          "step": "extraction2"
        }
      ]
    },
    {
      "name": "extraction2",
      "type": "EXTRACT",
      "config": {
        "extractorConfig": {
          "schema": {
            "type": "object",
            "properties": {
              "retailer": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Name of the retail store"
              },
              "transaction_date": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Date of transaction in DD/MM/YY format"
              },
              "transaction_time": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Time of transaction in HH:MM format"
              },
              "items": {
                "type": "array",
                "description": "List of items purchased",
                "items": {
                  "type": "object",
                  "properties": {
                    "name": {
                      "type": [
                        "string",
                        "null"
                      ],
                      "description": "Product name"
                    },
                    "price": {
                      "type": [
                        "number",
                        "null"
                      ],
                      "description": "Item price in GBP"
                    },
                    "quantity": {
                      "type": [
                        "number",
                        "null"
                      ],
                      "description": "Quantity purchased"
                    }
                  }
                }
              },
              "subtotal": {
                "type": [
                  "number",
                  "null"
                ],
                "description": "Subtotal amount in GBP"
              },
              "total": {
                "type": [
                  "number",
                  "null"
                ],
                "description": "Total amount paid in GBP"
              },
              "payment_method": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Payment method used"
              },
              "card_last_four": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Last four digits of card number"
              },
              "auth_code": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Authorization code for transaction"
              },
              "clubcard_number": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Loyalty card number (masked)"
              },
              "points_earned": {
                "type": [
                  "number",
                  "null"
                ],
                "description": "Loyalty points earned this visit"
              },
              "merchant_id": {
                "type": [
                  "string",
                  "null"
                ],
                "description": "Merchant identification number"
              }
            }
          }
        }
      }
    }
  ]
};

async function main() {
  console.log(`Deploying "${WORKFLOW.name}"…`);

  if (state.workflowId) {
    console.log(`✓ workflow already provisioned (${state.workflowId}) — updating steps`);
    await api("POST", `/workflows/${state.workflowId}`, { steps: WORKFLOW.steps });
  } else {
    // Reuse an existing workflow with the same name if one exists (e.g. a
    // previous run's state file was lost) instead of creating a duplicate.
    try {
      const list = await api("GET", `/workflows?name=${encodeURIComponent(WORKFLOW.name)}`);
      const items = (list.data ?? list.items ?? []) as Array<{ name?: string; id?: string }>;
      const existing = items.find((x) => x.name === WORKFLOW.name);
      if (existing?.id) {
        state.workflowId = existing.id; saveState();
        console.log(`✓ workflow "${WORKFLOW.name}" found in your account (${existing.id}) — updating steps`);
        await api("POST", `/workflows/${existing.id}`, { steps: WORKFLOW.steps });
      }
    } catch { /* lookup is best-effort; fall through to create */ }

    if (!state.workflowId) {
      const created = await api("POST", "/workflows", WORKFLOW);
      const wfId = created.id ?? created.workflow?.id;
      if (!wfId) throw new Error("Could not read created workflow id from response");
      state.workflowId = wfId; saveState();
      console.log(`+ created workflow (${wfId})`);
    }
  }

  // Deploy the current draft as a new version so the workflow is runnable —
  // best-effort: some accounts/plans may not require this explicit step.
  await api("POST", `/workflows/${state.workflowId}/versions`, {}).catch(() => {});

  console.log("\nDone. Run documents through it with:");
  console.log(`  POST ${API}/workflow_runs  { workflow: { id: "${state.workflowId}" }, file: { url: "https://…" } }`);
  console.log("Or open the workflow in the Extend dashboard to review and deploy it.");
}

main().catch((e) => { console.error(e.message ?? e); process.exit(1); });
import os
import json
import sys
from pathlib import Path
from typing import Any, Optional

from extend_ai import Extend

API_KEY = os.environ.get("EXTEND_API_KEY")
if not API_KEY:
    print("Set EXTEND_API_KEY first.", file=sys.stderr)
    sys.exit(1)

STATE_DIR = Path.cwd() / ".extend"
STATE_FILE = STATE_DIR / "receipt-extractor.json"

state: dict[str, Optional[str]] = {}
if STATE_FILE.exists():
    state = json.loads(STATE_FILE.read_text())


def save_state() -> None:
    STATE_DIR.mkdir(parents=True, exist_ok=True)
    STATE_FILE.write_text(json.dumps(state, indent=2))


WORKFLOW = {
    "name": "Receipt Processing Pipeline",
    "steps": [
        {
            "name": "startTrigger1",
            "type": "TRIGGER",
            "next": [{"step": "parse1"}],
        },
        {
            "name": "parse1",
            "type": "PARSE",
            "config": {
                "parseConfig": {
                    "blockOptions": {"text": {"agentic": {"enabled": True}}},
                    "chunkingStrategy": {"type": "document"},
                }
            },
            "next": [{"step": "extraction2"}],
        },
        {
            "name": "extraction2",
            "type": "EXTRACT",
            "config": {
                "extractorConfig": {
                    "schema": {
                        "type": "object",
                        "properties": {
                            "retailer": {
                                "type": ["string", "null"],
                                "description": "Name of the retail store",
                            },
                            "transaction_date": {
                                "type": ["string", "null"],
                                "description": "Date of transaction in DD/MM/YY format",
                            },
                            "transaction_time": {
                                "type": ["string", "null"],
                                "description": "Time of transaction in HH:MM format",
                            },
                            "items": {
                                "type": "array",
                                "description": "List of items purchased",
                                "items": {
                                    "type": "object",
                                    "properties": {
                                        "name": {
                                            "type": ["string", "null"],
                                            "description": "Product name",
                                        },
                                        "price": {
                                            "type": ["number", "null"],
                                            "description": "Item price in GBP",
                                        },
                                        "quantity": {
                                            "type": ["number", "null"],
                                            "description": "Quantity purchased",
                                        },
                                    },
                                },
                            },
                            "subtotal": {
                                "type": ["number", "null"],
                                "description": "Subtotal amount in GBP",
                            },
                            "total": {
                                "type": ["number", "null"],
                                "description": "Total amount paid in GBP",
                            },
                            "payment_method": {
                                "type": ["string", "null"],
                                "description": "Payment method used",
                            },
                            "card_last_four": {
                                "type": ["string", "null"],
                                "description": "Last four digits of card number",
                            },
                            "auth_code": {
                                "type": ["string", "null"],
                                "description": "Authorization code for transaction",
                            },
                            "clubcard_number": {
                                "type": ["string", "null"],
                                "description": "Loyalty card number (masked)",
                            },
                            "points_earned": {
                                "type": ["number", "null"],
                                "description": "Loyalty points earned this visit",
                            },
                            "merchant_id": {
                                "type": ["string", "null"],
                                "description": "Merchant identification number",
                            },
                        },
                    }
                }
            },
        },
    ],
}


def main() -> None:
    client = Extend(token=API_KEY)

    print(f'Deploying "{WORKFLOW["name"]}…"')

    if state.get("workflowId"):
        workflow_id = state["workflowId"]
        print(f"✓ workflow already provisioned ({workflow_id}) — updating steps")
        client.workflows.update(id=workflow_id, steps=WORKFLOW["steps"])
    else:
        existing_id: Optional[str] = None
        try:
            workflows_list = client.workflows.list(name=WORKFLOW["name"])
            items = workflows_list.data if hasattr(workflows_list, "data") else []
            for item in items:
                if item.name == WORKFLOW["name"]:
                    existing_id = item.id
                    break
        except Exception:
            pass

        if existing_id:
            state["workflowId"] = existing_id
            save_state()
            print(
                f'✓ workflow "{WORKFLOW["name"]}" found in your account ({existing_id}) — updating steps'
            )
            client.workflows.update(id=existing_id, steps=WORKFLOW["steps"])
        else:
            created = client.workflows.create(**WORKFLOW)
            workflow_id = created.id
            if not workflow_id:
                raise ValueError("Could not read created workflow id from response")
            state["workflowId"] = workflow_id
            save_state()
            print(f"+ created workflow ({workflow_id})")

    try:
        client.workflows.create_version(id=state["workflowId"])
    except Exception:
        pass

    print("\nDone. Run documents through it with:")
    print(
        f'  POST https://api.extend.ai/workflow_runs  {{ "workflow": {{ "id": "{state["workflowId"]}" }}, "file": {{ "url": "https://…" }} }}'
    )
    print("Or open the workflow in the Extend dashboard to review and deploy it.")


if __name__ == "__main__":
    try:
        main()
    except Exception as e:
        print(str(e), file=sys.stderr)
        sys.exit(1)
// This script uses the Extend REST API directly because Extend has no official Java SDK yet.
// Call the API with java.net.http.HttpClient and parse JSON manually.

import java.io.IOException;
import java.net.URI;
import java.net.URLEncoder;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.util.HashMap;
import java.util.Map;

public class ReceiptProvisioner {
  private static final String API = "https://api.extend.ai";
  private static final String VERSION = "2026-02-09";
  private static final String API_KEY = System.getenv("EXTEND_API_KEY");
  private static final Path STATE_DIR = Paths.get(System.getProperty("user.dir"), ".extend");
  private static final Path STATE_FILE = STATE_DIR.resolve("receipt-extractor.json");

  private static final HttpClient HTTP = HttpClient.newHttpClient();

  static class State {
    String workflowId;
  }

  private static State state = new State();

  public static void main(String[] args) throws Exception {
    if (API_KEY == null || API_KEY.isEmpty()) {
      System.err.println("Set EXTEND_API_KEY first.");
      System.exit(1);
    }

    loadState();

    String workflowName = "Receipt Processing Pipeline";
    System.out.println("Deploying \"" + workflowName + "\"…");

    String workflowDef = buildWorkflowJson();

    if (state.workflowId != null && !state.workflowId.isEmpty()) {
      System.out.println("✓ workflow already provisioned (" + state.workflowId + ") — updating steps");
      String stepsJson = extractStepsJson(workflowDef);
      api("POST", "/workflows/" + state.workflowId, "{\"steps\":" + stepsJson + "}");
    } else {
      // Try to find existing workflow by name
      try {
        String encoded = URLEncoder.encode(workflowName, StandardCharsets.UTF_8);
        String listResponse = api("GET", "/workflows?name=" + encoded, null);
        String existingId = findWorkflowIdByName(listResponse, workflowName);
        if (existingId != null) {
          state.workflowId = existingId;
          saveState();
          System.out.println("✓ workflow \"" + workflowName + "\" found in your account (" + existingId + ") — updating steps");
          String stepsJson = extractStepsJson(workflowDef);
          api("POST", "/workflows/" + existingId, "{\"steps\":" + stepsJson + "}");
        }
      } catch (Exception e) {
        // lookup is best-effort; fall through to create
      }

      if (state.workflowId == null || state.workflowId.isEmpty()) {
        String created = api("POST", "/workflows", workflowDef);
        String wfId = extractWorkflowId(created);
        if (wfId == null) {
          throw new Exception("Could not read created workflow id from response");
        }
        state.workflowId = wfId;
        saveState();
        System.out.println("+ created workflow (" + wfId + ")");
      }
    }

    // Deploy the current draft as a new version (best-effort)
    try {
      api("POST", "/workflows/" + state.workflowId + "/versions", "{}");
    } catch (Exception e) {
      // best-effort; some accounts/plans may not require this
    }

    System.out.println("\nDone. Run documents through it with:");
    System.out.println("  POST " + API + "/workflow_runs  { workflow: { id: \"" + state.workflowId + "\" }, file: { url: \"https://…\" } }");
    System.out.println("Or open the workflow in the Extend dashboard to review and deploy it.");
  }

  private static void loadState() throws IOException {
    if (Files.exists(STATE_FILE)) {
      String content = Files.readString(STATE_FILE);
      state.workflowId = extractJsonString(content, "workflowId");
    }
  }

  private static void saveState() throws IOException {
    Files.createDirectories(STATE_DIR);
    String json = "{\"workflowId\":\"" + escapeJson(state.workflowId) + "\"}";
    Files.writeString(STATE_FILE, json);
  }

  private static String api(String method, String pathName, String body) throws Exception {
    HttpRequest.Builder builder = HttpRequest.newBuilder()
        .uri(URI.create(API + pathName))
        .header("Authorization", "Bearer " + API_KEY)
        .header("x-extend-api-version", VERSION);

    if (body != null) {
      builder.header("Content-Type", "application/json")
          .method(method, HttpRequest.BodyPublishers.ofString(body));
    } else {
      builder.method(method, HttpRequest.BodyPublishers.noBody());
    }

    HttpRequest request = builder.build();
    HttpResponse<String> response = HTTP.send(request, HttpResponse.BodyHandlers.ofString());

    if (response.statusCode() < 200 || response.statusCode() >= 300) {
      String preview = response.body().length() > 300 ? response.body().substring(0, 300) : response.body();
      throw new Exception(method + " " + pathName + " failed (" + response.statusCode() + "): " + preview);
    }

    return response.body();
  }

  private static String buildWorkflowJson() {
    return "{"
        + "\"name\":\"Receipt Processing Pipeline\","
        + "\"steps\":["
        + "{"
        + "\"name\":\"startTrigger1\","
        + "\"type\":\"TRIGGER\","
        + "\"next\":[{\"step\":\"parse1\"}]"
        + "},"
        + "{"
        + "\"name\":\"parse1\","
        + "\"type\":\"PARSE\","
        + "\"config\":{"
        + "\"parseConfig\":{"
        + "\"blockOptions\":{\"text\":{\"agentic\":{\"enabled\":true}}},"
        + "\"chunkingStrategy\":{\"type\":\"document\"}"
        + "}"
        + "},"
        + "\"next\":[{\"step\":\"extraction2\"}]"
        + "},"
        + "{"
        + "\"name\":\"extraction2\","
        + "\"type\":\"EXTRACT\","
        + "\"config\":{"
        + "\"extractorConfig\":{"
        + "\"schema\":{"
        + "\"type\":\"object\","
        + "\"properties\":{"
        + "\"retailer\":{\"type\":[\"string\",\"null\"],\"description\":\"Name of the retail store\"},"
        + "\"transaction_date\":{\"type\":[\"string\",\"null\"],\"description\":\"Date of transaction in DD/MM/YY format\"},"
        + "\"transaction_time\":{\"type\":[\"string\",\"null\"],\"description\":\"Time of transaction in HH:MM format\"},"
        + "\"items\":{\"type\":\"array\",\"description\":\"List of items purchased\",\"items\":{\"type\":\"object\",\"properties\":{\"name\":{\"type\":[\"string\",\"null\"],\"description\":\"Product name\"},\"price\":{\"type\":[\"number\",\"null\"],\"description\":\"Item price in GBP\"},\"quantity\":{\"type\":[\"number\",\"null\"],\"description\":\"Quantity purchased\"}}}},"
        + "\"subtotal\":{\"type\":[\"number\",\"null\"],\"description\":\"Subtotal amount in GBP\"},"
        + "\"total\":{\"type\":[\"number\",\"null\"],\"description\":\"Total amount paid in GBP\"},"
        + "\"payment_method\":{\"type\":[\"string\",\"null\"],\"description\":\"Payment method used\"},"
        + "\"card_last_four\":{\"type\":[\"string\",\"null\"],\"description\":\"Last four digits of card number\"},"
        + "\"auth_code\":{\"type\":[\"string\",\"null\"],\"description\":\"Authorization code for transaction\"},"
        + "\"clubcard_number\":{\"type\":[\"string\",\"null\"],\"description\":\"Loyalty card number (masked)\"},"
        + "\"points_earned\":{\"type\":[\"number\",\"null\"],\"description\":\"Loyalty points earned this visit\"},"
        + "\"merchant_id\":{\"type\":[\"string\",\"null\"],\"description\":\"Merchant identification number\"}"
        + "}"
        + "}"
        + "}"
        + "}"
        + "}"
        + "]"
        + "}";
  }

  private static String extractStepsJson(String workflowJson) {
    int start = workflowJson.indexOf("\"steps\":");
    if (start == -1) return "[]";
    start += "\"steps\":".length();
    int depth = 0;
    int end = start;
    for (int i = start; i < workflowJson.length(); i++) {
      char c = workflowJson.charAt(i);
      if (c == '[') depth++;
      else if (c == ']') {
        depth--;
        if (depth == 0) {
          end = i + 1;
          break;
        }
      }
    }
    return workflowJson.substring(start, end);
  }

  private static String extractWorkflowId(String json) {
    String key = "\"id\":\"";
    int idx = json.indexOf(key);
    if (idx == -1) {
      key = "\"workflow\":{\"id\":\"";
      idx = json.indexOf(key);
      if (idx == -1) return null;
      idx += key.length();
    } else {
      idx += key.length();
    }
    int end = json.indexOf("\"", idx);
    if (end == -1) return null;
    return json.substring(idx, end);
  }

  private static String findWorkflowIdByName(String json, String targetName) {
    String searchKey = "\"name\":\"" + escapeJson(targetName) + "\"";
    int idx = json.indexOf(searchKey);
    if (idx == -1) return null;

    // Find the id field in the same object
    int objStart = json.lastIndexOf("{", idx);
    int objEnd = json.indexOf("}", idx);
    if (objStart == -1 || objEnd == -1) return null;

    String obj = json.substring(objStart, objEnd + 1);
    String idKey = "\"id\":\"";
    int idIdx = obj.indexOf(idKey);
    if (idIdx == -1) return null;
    idIdx += idKey.length();
    int idEnd = obj.indexOf("\"", idIdx);
    if (idEnd == -1) return null;
    return obj.substring(idIdx, idEnd);
  }

  private static String extractJsonString(String json, String key) {
    String searchKey = "\"" + key + "\":\"";
    int idx = json.indexOf(searchKey);
    if (idx == -1) return null;
    idx += searchKey.length();
    int end = json.indexOf("\"", idx);
    if (end == -1) return null;
    return json.substring(idx, end);
  }

  private static String escapeJson(String s) {
    if (s == null) return "";
    return s.replace("\\", "\\\\")
        .replace("\"", "\\\"")
        .replace("\n", "\\n")
        .replace("\r", "\\r")
        .replace("\t", "\\t");
  }
}
// This script uses the Extend REST API directly because Extend has no official Go SDK yet.
// Deploy the "Receipt" pipeline to YOUR Extend account.
//
// The workflow below is fully self-contained — every EXTRACT/CLASSIFY/SPLIT
// step carries its extractor/classifier/splitter config INLINE, so this is a
// single API call. No processors to create or wire up beforehand.
// Idempotent: the created workflow id is cached in .extend/receipt-extractor.json,
// so re-running updates the existing workflow instead of duplicating it.
//
// Usage:
//   export EXTEND_API_KEY=sk_...   (from https://dashboard.extend.ai → API Keys)
//   go run provision.go
//
// Generated by doc1 (template: receipt-extractor).

package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"io"
	"net/http"
	"net/url"
	"os"
	"path/filepath"
)

const (
	API     = "https://api.extend.ai"
	VERSION = "2026-02-09"
)

type State struct {
	WorkflowID string `json:"workflowId,omitempty"`
}

var (
	apiKey   string
	stateDir string
	stateFile string
	state    State
)

func init() {
	apiKey = os.Getenv("EXTEND_API_KEY")
	if apiKey == "" {
		fmt.Fprintf(os.Stderr, "Set EXTEND_API_KEY first.\n")
		os.Exit(1)
	}

	cwd, err := os.Getwd()
	if err != nil {
		fmt.Fprintf(os.Stderr, "Failed to get working directory: %v\n", err)
		os.Exit(1)
	}

	stateDir = filepath.Join(cwd, ".extend")
	stateFile = filepath.Join(stateDir, "receipt-extractor.json")

	// Load existing state if it exists
	if data, err := os.ReadFile(stateFile); err == nil {
		json.Unmarshal(data, &state)
	}
}

func saveState() error {
	if err := os.MkdirAll(stateDir, 0755); err != nil {
		return err
	}
	data, err := json.MarshalIndent(state, "", "  ")
	if err != nil {
		return err
	}
	return os.WriteFile(stateFile, data, 0644)
}

func apiCall(method, pathName string, body interface{}) (map[string]interface{}, error) {
	var reqBody io.Reader
	if body != nil {
		data, err := json.Marshal(body)
		if err != nil {
			return nil, err
		}
		reqBody = bytes.NewReader(data)
	}

	req, err := http.NewRequest(method, API+pathName, reqBody)
	if err != nil {
		return nil, err
	}

	req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", apiKey))
	req.Header.Set("x-extend-api-version", VERSION)
	if body != nil {
		req.Header.Set("Content-Type", "application/json")
	}

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return nil, err
	}
	defer resp.Body.Close()

	respBody, err := io.ReadAll(resp.Body)
	if err != nil {
		return nil, err
	}

	var data map[string]interface{}
	json.Unmarshal(respBody, &data)

	if resp.StatusCode >= 400 {
		respStr := string(respBody)
		if len(respStr) > 300 {
			respStr = respStr[:300]
		}
		return nil, fmt.Errorf("%s %s failed (%d): %s", method, pathName, resp.StatusCode, respStr)
	}

	return data, nil
}

var workflow = map[string]interface{}{
	"name": "Receipt Processing Pipeline",
	"steps": []map[string]interface{}{
		{
			"name": "startTrigger1",
			"type": "TRIGGER",
			"next": []map[string]interface{}{
				{
					"step": "parse1",
				},
			},
		},
		{
			"name": "parse1",
			"type": "PARSE",
			"config": map[string]interface{}{
				"parseConfig": map[string]interface{}{
					"blockOptions": map[string]interface{}{
						"text": map[string]interface{}{
							"agentic": map[string]interface{}{
								"enabled": true,
							},
						},
					},
					"chunkingStrategy": map[string]interface{}{
						"type": "document",
					},
				},
			},
			"next": []map[string]interface{}{
				{
					"step": "extraction2",
				},
			},
		},
		{
			"name": "extraction2",
			"type": "EXTRACT",
			"config": map[string]interface{}{
				"extractorConfig": map[string]interface{}{
					"schema": map[string]interface{}{
						"type": "object",
						"properties": map[string]interface{}{
							"retailer": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Name of the retail store",
							},
							"transaction_date": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Date of transaction in DD/MM/YY format",
							},
							"transaction_time": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Time of transaction in HH:MM format",
							},
							"items": map[string]interface{}{
								"type":        "array",
								"description": "List of items purchased",
								"items": map[string]interface{}{
									"type": "object",
									"properties": map[string]interface{}{
										"name": map[string]interface{}{
											"type":        []string{"string", "null"},
											"description": "Product name",
										},
										"price": map[string]interface{}{
											"type":        []string{"number", "null"},
											"description": "Item price in GBP",
										},
										"quantity": map[string]interface{}{
											"type":        []string{"number", "null"},
											"description": "Quantity purchased",
										},
									},
								},
							},
							"subtotal": map[string]interface{}{
								"type":        []string{"number", "null"},
								"description": "Subtotal amount in GBP",
							},
							"total": map[string]interface{}{
								"type":        []string{"number", "null"},
								"description": "Total amount paid in GBP",
							},
							"payment_method": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Payment method used",
							},
							"card_last_four": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Last four digits of card number",
							},
							"auth_code": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Authorization code for transaction",
							},
							"clubcard_number": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Loyalty card number (masked)",
							},
							"points_earned": map[string]interface{}{
								"type":        []string{"number", "null"},
								"description": "Loyalty points earned this visit",
							},
							"merchant_id": map[string]interface{}{
								"type":        []string{"string", "null"},
								"description": "Merchant identification number",
							},
						},
					},
				},
			},
		},
	},
}

func main() {
	fmt.Printf("Deploying \"%s\"…\n", workflow["name"])

	if state.WorkflowID != "" {
		fmt.Printf("✓ workflow already provisioned (%s) — updating steps\n", state.WorkflowID)
		steps := workflow["steps"]
		_, err := apiCall("POST", fmt.Sprintf("/workflows/%s", state.WorkflowID), map[string]interface{}{"steps": steps})
		if err != nil {
			fmt.Fprintf(os.Stderr, "%v\n", err)
			os.Exit(1)
		}
	} else {
		// Reuse an existing workflow with the same name if one exists
		found := false
		listPath := fmt.Sprintf("/workflows?name=%s", url.QueryEscape(workflow["name"].(string)))
		if list, err := apiCall("GET", listPath, nil); err == nil {
			var items []map[string]interface{}
			if data, ok := list["data"].([]interface{}); ok {
				for _, item := range data {
					items = append(items, item.(map[string]interface{}))
				}
			} else if data, ok := list["items"].([]interface{}); ok {
				for _, item := range data {
					items = append(items, item.(map[string]interface{}))
				}
			}

			for _, item := range items {
				if name, ok := item["name"].(string); ok && name == workflow["name"].(string) {
					if id, ok := item["id"].(string); ok {
						state.WorkflowID = id
						saveState()
						fmt.Printf("✓ workflow \"%s\" found in your account (%s) — updating steps\n", workflow["name"], id)
						steps := workflow["steps"]
						_, err := apiCall("POST", fmt.Sprintf("/workflows/%s", id), map[string]interface{}{"steps": steps})
						if err != nil {
							fmt.Fprintf(os.Stderr, "%v\n", err)
							os.Exit(1)
						}
						found = true
						break
					}
				}
			}
		}

		if !found {
			created, err := apiCall("POST", "/workflows", workflow)
			if err != nil {
				fmt.Fprintf(os.Stderr, "%v\n", err)
				os.Exit(1)
			}

			var wfID string
			if id, ok := created["id"].(string); ok {
				wfID = id
			} else if wf, ok := created["workflow"].(map[string]interface{}); ok {
				if id, ok := wf["id"].(string); ok {
					wfID = id
				}
			}

			if wfID == "" {
				fmt.Fprintf(os.Stderr, "Could not read created workflow id from response\n")
				os.Exit(1)
			}

			state.WorkflowID = wfID
			saveState()
			fmt.Printf("+ created workflow (%s)\n", wfID)
		}
	}

	// Deploy the current draft as a new version so the workflow is runnable
	apiCall("POST", fmt.Sprintf("/workflows/%s/versions", state.WorkflowID), map[string]interface{}{})

	fmt.Println("\nDone. Run documents through it with:")
	fmt.Printf("  POST %s/workflow_runs  { workflow: { id: \"%s\" }, file: { url: \"https://…\" } }\n", API, state.WorkflowID)
	fmt.Println("Or open the workflow in the Extend dashboard to review and deploy it.")
}

Frequently Asked Questions (FAQ)

Add explicit field descriptions like `"vendor_name": { description: "The business name printed at the top of the receipt, typically in bold or largest font" }` and `"total_amount": { description: "Final total including tax, often preceded by 'TOTAL:', 'AMOUNT DUE', or similar label" }`. Clear, specific descriptions are the #1 accuracy lever for receipts.
Use async (`parseRuns.createAndPoll`) for production—it handles large batches and complex receipts without timeout risk, though it adds some latency. Reserve sync parse only for real-time single-receipt flows under 10 pages.
Include currency symbol and locale context in your schema descriptions (e.g., `"total_amount": { description: "Total amount in EUR, e.g. '€49,99'" }`), and always set `type: ["string", "null"]` to capture the raw text without parsing—your downstream code can then normalize currency and locale.
Tags
ReceiptRetailTransactionLoyalty Program
About this template

This template captures retail receipt data including itemized product purchases with prices, payment method details (card authorization codes, merchant info), transaction totals, and loyalty program points. Designed for point-of-sale receipts.

Document formats
  • Images & Scans
  • PDF
Requirements
  • Long tables
  • Scanned documents