Faced with a large volume of PDF materials, manually opening each file to check information such as page count, title, author, subject, keywords, and creation date is not only time-consuming but also prone to omissions or errors. This article uses HeSoft Doc Batch Tool as an example to explain how to use the file information statistics feature to aggregate the path, name, extension, size, creation time, modification time, page count, and PDF metadata of multiple PDF files into an Excel spreadsheet in one go. It is suitable for scenarios such as document archiving, material inventory, project delivery checklist creation, teaching material statistics, and electronic file organization, helping users perform batch file information extraction with office software and reduce repetitive work.
When organizing PDF materials, many people encounter a seemingly simple but very time-consuming problem: there are dozens or hundreds of PDF files in a folder, and you need to compile the full path, file name, size, creation time, modification time, page count, and internal metadata such as title, author, subject, and keywords for each file into an Excel list. Manually opening each PDF to check its properties and then copying them into a spreadsheet is not only inefficient but also prone to issues like missed entries, incorrect data, and incomplete path copying.
This article introduces a method better suited for office scenarios: using the file information statistics function in HeSoft Doc Batch Tool to batch-process a large number of PDF files and compile the results into an Excel spreadsheet. Its core value is not modifying PDF content but centrally extracting information scattered across the file system and PDF properties, making it convenient for subsequent filtering, archiving, auditing, delivery, and management.
Applicable Scenarios: When to Batch-Process PDF Information
Batch-processing PDF file page counts and metadata is very common in many tasks. For example, administrative or archival staff need to organize electronic archive catalogs and must know the file path, file size, and creation time of each PDF; project teams need to create file lists when delivering materials, recording the name, page count, and version information of each PDF report; teachers or training institutions organizing courseware, handouts, and reading materials need to quickly understand the quantity and page count of a batch of PDF resources; legal, finance, research, and design positions also frequently need to conduct archival inventory checks on PDF documents.
Manual statistics might be acceptable for just three or five files. But when the number of PDFs increases to dozens or even more, manual operation becomes repetitive labor. Especially since PDF metadata isn't always directly displayed in the file explorer, viewing information like title, author, subject, keywords, creation date, application, and producer often requires opening file properties or using professional tools. Using office software for batch processing can combine these steps into a single operation.
It is important to note that this article discusses PDF file information statistics, not PDF content recognition or full-text extraction. That is to say, it primarily targets file lists, property statistics, page count statistics, and metadata compilation, with the final results presented in an Excel spreadsheet format, easy to copy, filter, sort, and further process.
Effect Preview: A Batch of Scattered PDF Files Before Processing
Before processing, the folder contains multiple PDF documents, such as ClassNotesPacket.pdf, CourseSummaryReport.pdf, ExamPreparationNotes.pdf, GrammarReviewHandout.pdf, etc. The file explorer can show the file name, modification date, type, and size, but this information is quite scattered and cannot directly provide the page count for each PDF or the PDF metadata fields. To create a material catalog from these files, you would also need to manually compile the full path and more properties.

From the state before processing, it's clear that the PDF files are numerous with similar names, making them suitable for unified batch statistics. For users who need to generate an archive catalog or submit an Excel list, relying solely on the folder list is far from enough, as the folder view cannot directly provide complete tabular results, nor is it convenient for filtering fields like page count, author, subject, or keywords.
Effect Preview: Automatically Generated Excel Statistics Sheet After Processing
After processing is complete, the PDF file information is compiled into an Excel spreadsheet. The table shows multiple types of fields, including path, name, name without extension, extension, size in bytes, size, name of the containing folder, path of the containing folder, creation time, modification time, page count, and PDF metadata-related information. The screenshot also displays columns for PDF metadata title, author, subject, keywords, creation date, application, and producer.

Such an Excel result is more complete than a manually created list and is also more suitable for subsequent processing. For example, you can sort by page count to quickly find the PDF with the most or fewest pages; filter by folder path to check which documents are in a specific directory; categorize by author, subject, or keywords; or save the list as a project delivery attachment, an electronic archive catalog, or an internal material ledger.
Operation Steps: Batch Export PDF Paths, Page Counts, and Metadata Using File Information Statistics
Step One: Enter the File Information Statistics Function in File Organization
After opening HeSoft Doc Batch Tool , select File Organization in the left function category. Locate File Information Statistics in the function card area. The function's description points to batch statistics for various files' names, paths, sizes, times, and metadata, perfectly suited for creating a PDF file list.

The purpose of this step is to enter the batch processing flow specifically designed for collecting file properties. Unlike checking properties of a single PDF, the File Information Statistics function targets a group of files, allowing multiple PDFs to be processed uniformly within the same task. After entering the function, subsequent steps follow a wizard-style guide to select files, configure statistics options, set the save location, and start processing.
Step Two: Add PDF Files or Import PDFs from a Folder
After entering the File Information Statistics interface, you first arrive at the record selection step. The top-right area of the interface provides two entry points: Add Files and Import Files from Folder. If you have a small number of PDF files, you can use Add Files to select specific PDFs; if the PDFs are all concentrated in a specific folder, it's more suitable to use Import Files from Folder to add the PDFs from that directory into the list all at once.

After files are imported, the list displays basic information such as sequence number, name, path, extension, creation time, and modification time. The screenshot shows that 20 records have been imported, and the file paths display as D:\test\folders\ followed by the corresponding PDF file name. Through this step, you can first confirm whether all the PDFs you need to process have been added. If you find mistakenly selected files, you can use the delete option in the operation column to remove them; if there are many files, you can also utilize filtering, sorting, and other entry points on the interface to assist in checking the list.
The expected result of this step is that a list of PDF records to be processed has been formed in the software. Only by confirming the list is correct before proceeding to the next setup step can you avoid missing statistics or over-processing.
Step Three: Check PDF File Detailed Information in Additional Information
After clicking next, you enter the processing options setup. The key area here is the Additional Information section. The interface shows options like Word File Detailed Information, Excel File Detailed Information, PPT File Detailed Information, PDF File Detailed Information, Text File Detailed Information, and Image File Detailed Information. Since this article aims to collect PDF page counts and PDF metadata, you need to check the box for PDF File Detailed Information.

This step is critical. Basic file information typically includes file name, path, extension, size, creation time, and modification time; PDF File Detailed Information is used to supplement PDF-specific fields, such as page count and PDF metadata like title, author, subject, keywords, creation date, application, and producer. Only by checking the corresponding detailed information will the exported Excel table contain these PDF-related columns.
If processing Word, Excel, PPT, image, or text files simultaneously, you can also check the corresponding detailed information options based on the actual file types. However, in this scenario, to obtain a cleaner PDF statistics table, you only need to focus on confirming that PDF File Detailed Information is checked.
Step Four: Set the Excel Save Location and Start Processing
After clicking next again, you enter the save location setup. Follow the on-screen wizard to choose a save location for the result file, which will store the final generated Excel statistics sheet. It is recommended to choose an easily accessible directory for the save location, such as the current project material directory, a temporary desktop folder, or a dedicated statistics results folder.
After setting the save location, proceed to the start processing step. Based on the previously imported PDF list and the checked PDF Detailed Information option, the software will batch-read each file's basic properties and PDF metadata and write these fields into the Excel table. This process eliminates the repetitive operations of opening each PDF individually, viewing properties, and manually copying paths and page counts.
The expected result of this step is an Excel file generated in the specified location. The higher the processing volume, the more obvious the efficiency advantage of batch processing becomes. For dozens of PDF files, manual statistics might take a considerable amount of time; using a batch processing tool can concentrate the work into import, confirmation, and export phases.
Step Five: Open the Excel File and Check the Results
After processing is complete, open the generated Excel table and check whether the column names and record count meet expectations. You can focus on checking the following categories of fields: First, file location fields, including path, name, name without extension, extension, containing folder name, and containing folder path; Second, file property fields, including size in bytes, size, creation time, and modification time; Third, PDF-specific fields, including page count, PDF metadata title, author, subject, keywords, creation date, application, and producer.
If some date columns in Excel display as hash marks, this is usually a display issue caused by insufficient column width. You can adjust the column width appropriately in Excel. As long as the cell's actual value exists, it can be viewed normally after adjusting the width.
Common Questions and Precautions
1. Why is some PDF metadata empty?
PDF metadata comes from the file's internal properties, and different PDFs are generated in different ways. Some PDFs exported from office software may contain information like title, author, subject, and keywords; others created by scanners, printers, or batch generation programs might only contain some fields or even have no metadata filled in. Therefore, it is normal for some metadata fields to be empty in the exported results and does not necessarily indicate a statistics failure.
2. What is the difference between page count and file size?
Page count indicates how many pages a PDF contains internally, suitable for estimating reading volume, printing volume, or material scale. File size indicates the storage space occupied by the PDF, typically presented in bytes or MB. The two fields serve different purposes: a file with many pages is not necessarily large, and a large PDF file does not necessarily have the most pages, because factors like image resolution, scan quality, and font embedding all affect file size.
3. Do PDFs need to be placed in the same folder before processing?
Not necessarily. The interface provides two methods: Add Files and Import Files from Folder. For operational convenience, it is recommended to gather the batch of PDFs needing statistics into a single project folder and then import them from the folder. This can reduce missed selections and make it easier to manage the statistical results by path later.
4. Can the exported Excel file be further edited?
Yes. After the statistical results are presented in Excel format, you can continue filtering, sorting, formatting, adding remark columns, or merging with other ledgers in Excel. For instance, you can add fields like archive number, person responsible, file status, and review comments to expand the PDF statistics list into a more comprehensive document management sheet.
5. Which PDF files are suitable for this process?
Any standard PDF file can be a target for statistics. Regardless of whether the file name is in English, numbers, or Chinese, basic properties can usually be read by the File Information Statistics function. For encrypted, damaged, or abnormally formatted PDFs, the specific information that can be read may be affected by the file status; it is recommended to check the corresponding rows in the result table after processing.
Summary: Delegate Repetitive PDF List Organization to Batch Processing Tools
Batch-processing PDF file paths, page counts, and metadata is essentially conducting a document asset inventory. While the manual method is intuitive, it suffers from low efficiency and high error rates when facing a large number of files, and it's not conducive to generating standardized results. With the file information statistics function of HeSoft Doc Batch Tool , you can compile the basic properties and detailed PDF information of multiple PDFs into an Excel table at once, significantly reducing repetitive labor.
If you are organizing project materials, teaching handouts, scanned archives, report files, or delivery documents, it is recommended to first consolidate the PDFs into a target folder, then follow the steps in this article to import files, check PDF File Detailed Information, set a save location, and start processing. The resulting Excel list can be directly used for archiving, checking, statistics, and delivery, making PDF file management clearer and more efficient.