Home >Backend Development >PHP Tutorial >How to read large files quickly with PHP_PHP Tutorial

How to read large files quickly with PHP_PHP Tutorial

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOriginal: 2016-07-13 10:17:23977browse

How to quickly read large files in PHP

In PHP, the fastest way to read files is to use some such as file, Functions such as file_get_contents can beautifully complete the functions we need with just a few lines of code. But when the file being operated is a relatively large file, these functions may be insufficient. The following will start with a requirement to explain the commonly used operating methods for reading large files.

Demand

There is an 800M log file with about 5 million lines. Use PHP to return the contents of the last few lines.

Implementation method

1. Directly use the file function to operate

Since the file function reads all the content into the memory at once , and PHP in order to prevent some poorly written programs from occupying too much memory and causing insufficient system memory and causing the server to crash, So by default, the maximum memory usage is limited to 16M, which is set through memory_limit = 16M in php.ini. If this value is set to -1, the memory usage Unrestricted.

The following is a piece of code that uses file to extract the last line of this file:

<?php
ini_set('memory_limit', '-1');
$file = 'access.log';
$data = file($file);
$line = $data[count($data) - 1];
echo $line;
?>

The entire code execution takes 116.9613 (s).

My machine has 2G of memory. When I press F5 to run, the system turns gray and only recovers after almost 20 minutes. It can be seen that the consequences of reading such a large file directly into the memory are serious, so As a last resort, memory_limit cannot be set too high, otherwise the only option is to call the computer room to reset the machine.

2. Directly call the Linux tail command to display the last few lines

In the Linux command line, you can directly use tail -n 10 access.log to easily display the last few lines of the log file. You can directly use PHP to call the tail command. The execution PHP code is as follows:

<?php
$file = 'access.log';
$file = escapeshellarg($file); // 对命令行参数进行安全转义
$line = `tail -n 1 $file`;
echo $line;
?>

The entire code execution takes 0.0034 (s)

3. Directly use PHP’s fseek to perform file operations

This method is the most common method. It does not need to read all the contents of the file, but operates directly through the pointer, so the efficiency is quite efficient. When using fseek to operate files, there are many different methods, and the efficiency may be slightly different. The following are two commonly used methods:

Method 1

First find the last EOF of the file through fseek, then find the starting position of the last line, take the data of this line, then find the starting position of the next line, then take the position of this line, and so on, until found $num rows.

#The implementation code is as follows

<?php
$fp = fopen($file, "r");
$line = 10;
$pos = -2;
$t = " ";
$data = "";
while ($line > 0)
{
	while ($t != "\n")
	{
		fseek($fp, $pos, SEEK_END);
		$t = fgetc($fp);
		$pos--;
	}
	$t = " ";
	$data .= fgets($fp);
	$line--;
}
fclose($fp);
echo $data
?>

The entire code execution takes 0.0095 (s)

Method 2

Still use fseek to read from the end of the file, but this time it is not reading bit by bit, but reading piece by piece. Every time a piece of data is read, the read data is placed in a buf. , and then use the number of newline characters (n) to determine whether the last $num rows of data have been read.

#The implementation code is as follows

<?php
$fp = fopen($file, "r");
$num = 10;
$chunk = 4096;
$fs = sprintf("%u", filesize($file));
$max = (intval($fs) == PHP_INT_MAX) ? PHP_INT_MAX : filesize($file);
for ($len = 0; $len < $max; $len += $chunk)
{
	$seekSize = ($max - $len > $chunk) ? $chunk : $max - $len;
	fseek($fp, ($len + $seekSize) * -1, SEEK_END);
	$readData = fread($fp, $seekSize) . $readData;
	if (substr_count($readData, "\n") >= $num + 1)
	{
		preg_match("!(.*?\n){" . ($num) . "}$!", $readData, $match);
		$data = $match[0];
		break;
	}
}
fclose($fp);
echo $data;
?>

The entire code execution takes 0.0009(s).

Method 3

<?php
function tail($fp, $n, $base = 5)
{
	assert($n > 0);
	$pos = $n + 1;
	$lines = array();
	while (count($lines) <= $n)
	{
		try
		{
			fseek($fp, -$pos, SEEK_END);
		}
		catch (Exception $e)
		{
			fseek(0);
			break;
		}
		$pos *= $base;
		while (!feof($fp))
		{
			array_unshift($lines, fgets($fp));
		}
	}

	return array_slice($lines, 0, $n);
}

var_dump(tail(fopen("access.log", "r+"), 10));
?>

The entire code execution takes 0.0003(s)

Articles you may be interested in

php function to read the directory and list the files in the directory
Summary of php reading xml files
PHP uses Curl Functions to implement multi-threaded crawling of web pages and downloading files
php error_log() writes error information to a file
PHP’s method of serializing variables competes with four methods of serializing variables in PHP
Use PHP’s GZip compression function to compress website JS and CSS files to speed up website access
PHP Delete the directory and all files in the directory
How to solve the problem of concurrent reading and writing file conflicts in php

Statement：

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Previous article：Problems with accessing the symfony framework under China Mobile cmwap network, symfonycmwap_PHP tutorialNext article：Problems with accessing the symfony framework under China Mobile cmwap network, symfonycmwap_PHP tutorial

See more

How to read large files quickly with PHP_PHP Tutorial

How to quickly read large files in PHP

Articles you may be interested in

Related articles