Home >Backend Development >PHP Tutorial >PHP Master | Extract Objects from an Access Database with PHP, Part 2

PHP Master | Extract Objects from an Access Database with PHP, Part 2

William Shakespeare
William ShakespeareOriginal
2025-02-24 10:45:10300browse

This article demonstrates how to extract embedded PDF and image files from legacy Microsoft Access databases using PHP. Part 1 covered extracting packaged objects; this part focuses on PDFs and common image formats (BMP, GIF, JPEG, PNG). These files, while diverse, share a common OLE container structure: a variable-length header and trailer. We'll leverage this structure for extraction.

Key Concepts:

  • PDF Extraction: PHP's strpos() and substr() functions pinpoint and extract PDFs by identifying the hexadecimal sequences %PDF (25504446) and %%EOF (2525454F46).
  • Image Extraction (BMP, GIF, JPEG, PNG): Similar techniques are used, adapting the start and end delimiters for each image type.
  • Handling Unknown OLE Types: A new function, extractUnknown(), saves unidentified OLE objects for later analysis, enhancing the script's robustness.
  • Enhanced Switch Statement: The original switch statement is improved to handle a wider range of OLE object types.

Extracting Adobe Acrobat Documents (PDFs)

The example database contains a PDF in record 13. Inspecting the OLE field's initial bytes reveals the PDF's presence but lacks metadata like filename or size. However, the consistent %PDF and %%EOF markers in all PDFs allow for reliable extraction. The PHP script searches for these hexadecimal sequences to determine the start and end points, enabling extraction using substr().

PHP Master | Extract Objects from an Access Database with PHP, Part 2

PHP Master | Extract Objects from an Access Database with PHP, Part 2

Handling Other Object Types

The improved PHP script includes extractUnknown() to handle and save unknown OLE types (using the record ID as the filename) for later examination. This is crucial for identifying embedded images.

<code class="language-php"><?php
function extractUnknown($id, $data) {
    file_put_contents($id . ".txt", hex2bin($data));
}
?></code>

Extracting Popular Image Types

Image type identification within the OLE header varies depending on the originating software and file associations. The extractUnknown() function helps catalog these types. We'll focus on BMP, GIF, JPEG, and PNG. GIF, JPEG, and PNG extraction mirrors the PDF method, changing only the delimiters:

PHP Master | Extract Objects from an Access Database with PHP, Part 2

BMP extraction is slightly different. The start is easily found (BM), but the end requires calculating the size (from the header) and converting it to big-endian format before using it to extract the data.

PHP Master | Extract Objects from an Access Database with PHP, Part 2

Complete PHP Script (Partial)

The following is a snippet of the updated PHP script. The functions for extracting GIF, JPEG, and PNG are omitted for brevity but follow the same pattern as PDF and BMP extraction.

<code class="language-php"><?php
function extractUnknown($id, $data) {
    file_put_contents($id . ".txt", hex2bin($data));
}
?></code>

The complete, updated script (including the omitted functions) is available on GitHub (links to part-1 and part-2 branches). This improved script offers a more comprehensive solution for extracting various OLE object types from Access databases. This two-part series provides valuable tools for migrating away from legacy Access databases.

(FAQs section omitted for brevity, but could be re-written in a similar paraphrased style to the rest of the output.)

The above is the detailed content of PHP Master | Extract Objects from an Access Database with PHP, Part 2. For more information, please follow other related articles on the PHP Chinese website!

Statement:
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn