Home >Backend Development >PHP Tutorial >PHP Master | Extract Objects from an Access Database with PHP, Part 2
This article demonstrates how to extract embedded PDF and image files from legacy Microsoft Access databases using PHP. Part 1 covered extracting packaged objects; this part focuses on PDFs and common image formats (BMP, GIF, JPEG, PNG). These files, while diverse, share a common OLE container structure: a variable-length header and trailer. We'll leverage this structure for extraction.
Key Concepts:
strpos()
and substr()
functions pinpoint and extract PDFs by identifying the hexadecimal sequences %PDF
(25504446) and %%EOF
(2525454F46).extractUnknown()
, saves unidentified OLE objects for later analysis, enhancing the script's robustness.Extracting Adobe Acrobat Documents (PDFs)
The example database contains a PDF in record 13. Inspecting the OLE field's initial bytes reveals the PDF's presence but lacks metadata like filename or size. However, the consistent %PDF
and %%EOF
markers in all PDFs allow for reliable extraction. The PHP script searches for these hexadecimal sequences to determine the start and end points, enabling extraction using substr()
.
Handling Other Object Types
The improved PHP script includes extractUnknown()
to handle and save unknown OLE types (using the record ID as the filename) for later examination. This is crucial for identifying embedded images.
<code class="language-php"><?php function extractUnknown($id, $data) { file_put_contents($id . ".txt", hex2bin($data)); } ?></code>
Extracting Popular Image Types
Image type identification within the OLE header varies depending on the originating software and file associations. The extractUnknown()
function helps catalog these types. We'll focus on BMP, GIF, JPEG, and PNG. GIF, JPEG, and PNG extraction mirrors the PDF method, changing only the delimiters:
BMP extraction is slightly different. The start is easily found (BM
), but the end requires calculating the size (from the header) and converting it to big-endian format before using it to extract the data.
Complete PHP Script (Partial)
The following is a snippet of the updated PHP script. The functions for extracting GIF, JPEG, and PNG are omitted for brevity but follow the same pattern as PDF and BMP extraction.
<code class="language-php"><?php function extractUnknown($id, $data) { file_put_contents($id . ".txt", hex2bin($data)); } ?></code>
The complete, updated script (including the omitted functions) is available on GitHub (links to part-1 and part-2 branches). This improved script offers a more comprehensive solution for extracting various OLE object types from Access databases. This two-part series provides valuable tools for migrating away from legacy Access databases.
(FAQs section omitted for brevity, but could be re-written in a similar paraphrased style to the rest of the output.)
The above is the detailed content of PHP Master | Extract Objects from an Access Database with PHP, Part 2. For more information, please follow other related articles on the PHP Chinese website!