Official Node.js SDK for Invoice Data Extraction. Uploads your files, submits the extraction, waits for it to finish and hands you the rows as data, or a spreadsheet, in a few lines of code.
- Node.js 18 or later
- ESM only
npm install @invoicedataextraction/sdkThis package is ESM only. Your project's package.json must include "type": "module" (or use .mjs file extensions). TypeScript declarations are included.
import InvoiceDataExtraction from "@invoicedataextraction/sdk";
const client = new InvoiceDataExtraction({
api_key: process.env.INVOICE_DATA_EXTRACTION_API_KEY,
});
const result = await client.extract({
folder_path: "./invoices",
prompt: "Extract invoice number, date, vendor name, and total amount",
output_structure: "per_invoice",
json_typed_values: true,
console_output: true, // remove to disable console logging
});
if (result.status === "completed") {
for await (const row of client.iterateResults({ extraction_id: result.extraction_id })) {
console.log(row); // one object per extracted row, keyed by your output columns
}
}extract(...) uploads your files (pass a folder_path or a list of files), submits the extraction, waits until it finishes and returns the result: the final status response from the API, for a completed, failed or cancelled extraction. iterateResults(...) then reads the extracted rows straight from the API, with amounts as numbers and empty cells as null because the extraction was submitted with json_typed_values; getResults(...) reads one page when you want to manage paging yourself. Check result.pages.failed_count to verify that all uploaded pages were processed, and result.review_needed.count for rows that need a human's check before you rely on the data.
To get a spreadsheet, add download: { formats: ["xlsx"], output_path: "./output" } to the call and the file is saved when the extraction completes, or call downloadOutput(...) later.
Generate an API key from your dashboard. Every account includes 50 free pages per month. Additional credits can be purchased on a pay-as-you-go basis with no subscription needed.
If you need control over individual steps, for example uploading files in one part of your system and extracting in another, use the lower-level methods:
const upload = await client.uploadFiles({
files: ["./invoice1.pdf", "./invoice2.pdf"],
console_output: true,
});
const submitted = await client.submitExtraction({
upload_session_id: upload.upload_session_id,
file_ids: upload.file_ids,
prompt: "Extract invoice number and total",
output_structure: "per_invoice",
json_typed_values: true,
});
const result = await client.waitForExtractionToFinish({
extraction_id: submitted.extraction_id,
console_output: true,
});
for await (const row of client.iterateResults({ extraction_id: submitted.extraction_id })) {
console.log(row);
}Beside json_typed_values, extract(...) and submitExtraction(...) take output_language, review_needed_fill_color, affected_field_fill_color and send_completion_email, each applying to that extraction only. Without json_typed_values, every value in the JSON output and in the rows is a string.
Set ask_questions: true and the extraction can stop to ask when the documents leave something unsettled, instead of deciding on its own. Pass on_questions and the SDK calls it as on_questions(questions, status), sends back the answers it returns and carries on to the result; without it, extract(...) returns status: "input_required" with the questions for you to answer with answerQuestions(...). The questions, the answer forms and the deadline are in the API reference.
const result = await client.extract({
folder_path: "./invoices",
prompt: "Extract invoice number, date, vendor name, and total amount",
output_structure: "per_invoice",
json_typed_values: true,
ask_questions: true,
on_questions: async (questions) =>
questions.map((question) => ({
question_id: question.question_id,
accept_recommended: true, // or { choice_id: "b" }, or { text: "DD/MM/YYYY" }
})),
});
// Stop an extraction that is still queued or processing
await client.cancelExtraction({ extraction_id });// One-page browse with filters
const page = await client.listExtractions({
status: "completed",
limit: 50,
});
// Auto-paginating iterator over every matching extraction
for await (const extraction of client.iterateExtractions({ status: "completed" })) {
console.log(extraction.extraction_id, extraction.task_name);
}
// Full record (the original prompt, options, full pages, full failure error, etc.)
const { extraction } = await client.getExtraction({ extraction_id: "..." });Team admins can pass scope: "team" on listing methods to browse team-visible extractions, or scope: "own" to force own-only listing.
SDK methods reject with a normal JavaScript Error. The structured error body is on error.body:
try {
await client.extract({ /* ... */ });
} catch (error) {
console.log(error.body.error.code); // e.g. "INVALID_INPUT"
console.log(error.body.error.message); // Human-readable message
console.log(error.body.error.retryable);
}When an extraction task itself reaches a terminal state, extract(...) returns that response rather than throwing: check result.status for "completed", "failed", or "cancelled" for tasks stopped from the web app or with cancelExtraction(...).
- Node SDK docs: full method reference, parameters, return shapes, and examples
- REST API docs: endpoint-level documentation for direct HTTP integration
- Dashboard: manage API keys and view extraction results
MIT