How to batch extract PDF page counts and metadata and generate an Excel file list


Translation:EnglishFrançaisDeutschEspañol日本語한국어,Update Time:2026-08-02 06:31:08

Disclaimer: All images, text, and video content on the website are for reference only and may not be the latest, correct, or accurate. In case of any dispute, please refer to the actual experience effect!

When creating a PDF file inventory, common fields go beyond just file name and size, and also include the full path, creation date, modification date, page count, and metadata such as PDF title, author, subject, and keywords. This article uses HeSoft Doc Batch Tool as an example to explain how to batch import PDFs using the file information statistics feature, select detailed PDF file information, and generate an Excel summary table, helping users quickly complete asset inventory, archive catalog organization, and document delivery checks.

When many people organize PDF files, their first instinct is to open Excel, then browse folders while manually entering file names, paths, sizes, and dates. This approach is fine for just a few files, but when dealing with dozens or hundreds of documents, manual entry becomes extremely inefficient. Even more troublesome, information like PDF page counts, titles, authors, subjects, keywords, creation dates, applications, and producer programs often cannot be fully viewed directly in folder lists—you need to open the file or check its properties to confirm.

This article introduces a method better suited for office scenarios: using the "File Information Statistics" feature in HeSoft Doc Batch Tool to batch-extract basic attributes and detailed PDF information from numerous PDF files and compile them into an Excel spreadsheet. The entire process follows a wizard-based approach, from selecting files and setting processing options to saving and starting the process, making it ideal for users who need to quickly create PDF file lists, PDF page count tables, or PDF metadata ledgers.

Applicable Scenarios: From PDF Directory Organization to Metadata Review

The need to batch-extract PDF page counts and metadata is quite common in office settings. It often arises in the following types of work:

  • Archive Catalog Creation: Organizing scanned documents, electronic contracts, policy documents, and report files into an Excel directory for easy archiving and retrieval.
  • Teaching Material Statistics: Counting pages in PDFs like courseware, handouts, study materials, practice manuals, and reading materials to assess resource volume.
  • Project Document Delivery: Generating a file list before submitting project materials to check if PDFs are complete, consistently named, and have correct paths.
  • File Source Verification: Using metadata fields like author, application, and producer program in PDFs to help determine file origin and generation method.
  • Batch File Management: Filtering abnormal documents by fields like size, modification time, and containing folder to discover duplicate or outdated materials.

In these scenarios, Excel is the most commonly used summarization tool because it's easy to sort, filter, search, and share. The batch processing capabilities of office software can consolidate file information that would otherwise need manual verification one by one.

Effect Preview: Changes Before and After Batch Statistics

Before Processing: PDFs Scattered in Folders, Only Limited Attributes Visible

The screenshot below before processing shows a folder containing multiple PDFs. File names include ClassNotesPacket.pdf, CourseSummaryReport.pdf, ExamPreparationNotes.pdf, KnowledgeReviewNotes.pdf, etc. From Windows File Explorer, you can see basic information like file names, modification dates, types, and sizes, but this information is incomplete. For instance, you cannot directly see how many pages each PDF has, nor can you batch-view PDF titles, authors, subjects, or keywords.

image-Batch extract PDF page count,bulk export PDF metadata,PDF file list to Excel

If you need to compile these PDFs into a formal ledger, relying solely on the folder view is insufficient. Users typically also need the full path for locating original files; creation and modification times for judging file versions; page counts for estimating material scale; and metadata for checking document standardization.

After Processing: A Structured List in Excel with One Row per PDF

After processing, the PDF file information is exported to Excel. The screenshot shows that the table is no longer just a list of file names but contains very comprehensive fields: path, name, name (without extension), extension, size (bytes), size, containing folder name, containing folder path, creation time, modification time, page count, and PDF metadata like title, author, subject, keywords, creation date, application, producer program, etc.

image-Batch extract PDF page count,bulk export PDF metadata,PDF file list to Excel

This Excel result is much more suitable for subsequent office processing. For example, to see which PDFs have more than 10 pages, you can sort by the "Page Count" column; to check if a batch of PDFs was generated by the same application, you can filter by the "PDF - Metadata - Application" column; to organize source materials by author, you can filter or group by the "PDF - Metadata - Author" column.

Steps: Batch Export PDF Page Counts, Paths, and Metadata

Step 1: Open the File Information Statistics Entry Point

After launching HeSoft Doc Batch Tool , select "File Management" in the left navigation. In the function card area, you will see "File Information Statistics," described as batch statistics for various file names, paths, sizes, times, metadata, and other information. Since we want to extract PDF file information, click this function to enter.

image-Batch extract PDF page count,bulk export PDF metadata,PDF file list to Excel

The purpose of this step is to select the correct batch processing tool. Do not choose "Folder Information Statistics," as the goal of this article is to count specific PDF files, not the folders themselves. After entering "File Information Statistics," the software guides you through a step-by-step wizard to help with importing, setting options, specifying a save location, and starting the process.

Step 2: Import the PDF Files to be Counted

After entering the function page, you are currently at Step 1, "Select records to process." The top right of the page provides two common entry points: "Add Files" and "Import Files from Folder." If the PDFs are relatively concentrated, it's recommended to use "Import Files from Folder," which can load all files from the target folder into the list at once. If you only need to count a few PDFs, you can use "Add Files."

image-Batch extract PDF page count,bulk export PDF metadata,PDF file list to Excel

After importing is complete, the table will list file sequence numbers, names, paths, extensions, creation times, modification times, and other information. The bottom of the screenshot shows a record count of 20, indicating that 20 PDF files have been added for this task. It's advisable to do a quick check at this point: Are the file extensions .pdf? Do the paths belong to the target folder? Does the file count match expectations?

If you find any mistakenly selected files, you can delete individual records from the operation column; if you need to re-import, you can use "Clear." When dealing with many files, the "Filter" and "Sort" functions provided on the page can also help quickly review the list. Once confirmed correct, click "Next" at the bottom.

Step 3: Select PDF File Detailed Information in Additional Information

Upon entering Step 2, "Set processing options," you'll see the "Additional Information" area. This section lists detailed information options for various file types, including Word file details, Excel file details, PPT file details, PDF file details, text file details, and image file details.

image-Batch extract PDF page count,bulk export PDF metadata,PDF file list to Excel

For this task, which involves batch-extracting PDF page counts and metadata, you need to check "PDF File Detailed Information." This is the key setting for generating a complete PDF list. Once checked, the software will read PDF-related detailed fields in addition to basic file attributes. The columns like "Page Count," "PDF - Metadata - Title," "PDF - Metadata - Author," "PDF - Metadata - Subject," "PDF - Metadata - Keywords" that appear in the final Excel output are precisely a result of selecting this option.

If you are also processing Word, Excel, PPT, images, or other file types simultaneously, you can check their corresponding detailed information based on actual needs. However, for a clearer result table, it's recommended to only import PDFs this time, or at least ensure that PDFs are the primary focus files for analysis.

Step 4: Set the Save Location for the Statistics Results

In the top workflow, Step 3 is "Set save location." After clicking "Next" to enter, follow the on-screen prompts to choose where the result file should be saved. It's advisable to save the output result to an easily locatable place, such as the current PDF project directory, the archival directory, or a dedicated statistics results folder.

The purpose of setting the save location is to avoid being unable to find the exported Excel file after processing is complete. For users who need to count different batches of PDFs multiple times, you can also include dates, project names, or material categories in the saved file name to easily distinguish summary sheets from different batches.

Step 5: Start Processing and Review the Excel Result

After completing the save location settings, proceed to Step 4, "Start processing." The software will read PDF file information one by one according to the task list and automatically organize it into a table result. Compared to manual statistics, the main advantages of batch processing are stability, consistency, and repeatability: the same set of fields is used for the same batch of files, avoiding issues like inconsistent column names, non-uniform date formats, and incomplete path copying common in manual entry.

Once processing is complete, open the Excel result. It's recommended to prioritize checking whether the following fields meet your requirements: Can the full path locate the original file? Has the page count been generated? Do PDF metadata columns like Title, Author, Subject, Keywords exist? Can creation time and modification time be used for version judgment? Can file size be used to check for anomalous files?

Common Questions and Notes

1. If I only want to count PDF page counts, do I need to export all metadata?

Judging from the screenshot results, checking PDF file detailed information outputs page counts and multiple types of PDF metadata fields. If you only care about page counts, you can keep the "Page Count," necessary path, and file name columns in Excel after generation, and delete or hide the unneeded metadata columns. This way, you can leverage the software's complete extraction while keeping the final table concise.

2. What is the difference between the PDF file name and the name in Excel?

The result table typically includes both "Name" and "Name (without extension)." The former includes the .pdf extension, which is convenient for cross-referencing with the original file; the latter removes the extension, which is suitable for subsequent matching and organization by title, number, or material name.

3. Why is the full path field needed?

When PDFs with the same name exist in multiple folders, the file name alone might not be enough to determine which specific file it is. The full path clarifies the file's location, facilitating backtracking, opening the original file, or performing subsequent migrations.

4. Does incomplete metadata mean the software failed to read it?

Not necessarily. Whether PDF metadata is complete depends on the file itself. Many PDFs are generated without filling in author, subject, or keywords, so it's normal for corresponding fields to be empty. The software can batch-extract existing information but cannot generate metadata that does not exist within the file.

5. How to reduce the chance of errors when dealing with a large number of files?

It's recommended to import files in batches by folder or project, confirm the record count and paths first, and then check "PDF File Detailed Information." After processing, use Excel to filter for null values, unusual page counts, or anomalous times for a secondary check. This maintains processing efficiency while also making it easy to pinpoint problematic files.

Summary: Using Batch Processing to Complete PDF Page Count and Metadata Statistics

The core of creating a PDF file list isn't simply listing file names, but rather uniformly consolidating various information scattered across the file system and within the PDFs themselves. Through its "File Information Statistics" feature, HeSoft Doc Batch Tool connects the steps of importing files, selecting PDF detailed information, setting a save location, and starting processing, allowing users to quickly obtain a PDF information sheet in Excel format.

If you are currently working on PDF archiving, document delivery, page counting, or metadata review, you can follow the process outlined in this article: first enter File Information Statistics within File Management, then import PDF files, check "PDF File Detailed Information," and finally export to Excel. Compared to opening PDFs one by one to view, this method is better suited for large-volume file processing in office scenarios, significantly reducing repetitive work and improving statistical accuracy and delivery efficiency.


Keyword:Batch extract PDF page count , bulk export PDF metadata , PDF file list to Excel
Creation Time:2026-08-02 06:30:52

Disclaimer: All images, text, and video content on the website are for reference only and may not be the latest, correct, or accurate. In case of any dispute, please refer to the actual experience effect!

Related Articles