pdfextract API
    Preparing search index...

    Interface PdfDocument

    An open PDF handle. Operations are serialized per document; release it with close in a finally block.

    interface PdfDocument {
        pageCount: number;
        close(): Promise<void>;
        extractImage(
            id: string,
            options?: ExtractImageOptions,
        ): Promise<ExtractedImage>;
        getImages(options?: GetImagesOptions): Promise<readonly EmbeddedImage[]>;
        getStructuredText(options?: StructuredTextOptions): Promise<StructuredText>;
    }
    Index
    • Idempotently cancel queued work and join engine cleanup. Returned data remains valid; the OCR provider stays open.

      Returns Promise<void>

    pageCount: number

    Total pages in the input PDF, independent of per-operation page selection.