URL crawler-PHP Tutorial-php.cn

Home

Backend Development

PHP Tutorial

URL crawler

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWB

Jul 25, 2016 am 08:48 AM

If you need points-free download of csdn, point-free download of pudn, or point-free 51cto, please go to http://www.itziy.com/
and execute it from the command line. Direct php call will display the usage method
Function description
1. Support agent
2. Supports setting the number of recursive checks
3. Supports output type control and check content control

Function:
主要代替肉眼尽量多的抓取可能的请求包及url地址等，方便渗透测试

error_reporting(E_ERROR | E_WARNING | E_PARSE);
ini_set('memory_limit','1024M');
set_time_limit(0);
define('CHECK_A_TAG', false);
define('CHECK_JS_TAG', true);
define('CHECK_URL', true);
define('SAVE_ERROR', true);
$checkArr = array(
'$.load',
'.ajax',
'$.post',
'$.get',
'.getJSON'
);
if ($argc die(showerror('sorry, parameter error', array('example: php debug.php url num filename header proxy', 'detail information:', 'url: target url address which you want to check it', 'num: The number of pages of recursive,default 3', 'filename: output filename default name ret.txt', 'header: The request header file default null', 'proxy: if you want to use proxy set it here default no use proxy')));
if (!check_extension())
die(showerror('extension curl not support', 'please open php curl extension support'));
//global variable
$url = trim($argv[1]);
if (stripos($url, 'http') === false)
$url = 'http://'.$url;
$num = isset($argv[2]) ? intval($argv[2]) : 3;
$output = isset($argv[3]) ? trim(str_replace("\", '/', $argv[3])) : str_replace("\", '/', dirname(__FILE__)).'/ret.txt';
$header = null;
$proxy = null;
$host = null;
if (isset($argv[4]))
{
$header = trim(str_replace("\", '/', $argv[4]));
if (file_exists($header))
$header = array_filter(explode("n", str_replace("r", '', file_get_contents($header))));
else
{
$file = str_replace("\", '/', dirname(__FILE__)).'/'.$header;
if (file_exists($file))
$header = array_filter(explode("n", str_replace("r", '', file_get_contents($file))));
else
$header = null;
}
}
if (isset($argv[5]))
$proxy = trim($argv[5]);
if (!is_array($header) || empty($header))
$header = null;
$result = check_valid_url($url);
$outputArr = array();
if (!empty($result))
{
$result = str_replace("r", '', $result);
$result = str_replace("n", '', $result);
$tmpArr = parse_url($url);
if (!isset($tmpArr['host']))
die(showerror('parse url error', 'can not get host form url: '.$url));
$host = $tmpArr['host'];
if (stripos($host, 'http') === false)
$host = 'http://'.$host;
unset($tmpArr);
//check for current page
if (!isset($outputArr[md5($url)]))
{
$outputArr[md5($url)] = $url;
file_put_contents($output, $url."n", FILE_APPEND);
echo 'url: ',$url,' find ajax require so save it',PHP_EOL;
}
work($result);
}
echo 'run finish',PHP_EOL;
function work($result, $reverse = false)
{
global $num, $host, $outputArr, $checkArr, $output;
if (!$result)
return;
$result = str_replace("r", '', $result);
$result = str_replace("n", '', $result);
while ($num > 0)
{
echo 'remain: ',$num,' now start to check for url address',PHP_EOL,PHP_EOL;
preg_match_all('//i', $result, $match);
if (CHECK_A_TAG && isset($match[2]) && !empty($match[2]))
{
foreach ($match[2] as $mc)
{
$mc = trim($mc);
if ($mc == '#')
continue;
if (stripos($mc, 'http') === false)
$mc = $host.$mc;
if (($ret = check_valid_url($mc)))
{
if (!isset($outputArr[md5($mc)]))
{
$outputArr[md5($mc)] = $mc;
file_put_contents($output, $mc."n", FILE_APPEND);
echo 'url: ',$mc,' find ajax require so save it',PHP_EOL;
}
}
}
}
//check for page url
echo 'remain: ',$num,' now start to check for page url',PHP_EOL,PHP_EOL;
preg_match_all('/(https?|ftp|mms)://([A-z0-9]+[_-]?[A-z0-9]+.)*[A-z0-9]+-?[A-z0-9]+.[A-z]{2,}(/.*)*/?/i', $result, $match);
if (CHECK_URL && isset($match[2]) && !empty($match[2]))
{
foreach ($match[2] as $mc)
{
$mc = trim($mc);
if ($mc == '#')
continue;
if (stripos($mc, 'http') === false)
$mc = $host.$mc;
if (($ret = check_valid_url($mc)))
{
if (!isset($outputArr[md5($mc)]))
{
$outputArr[md5($mc)] = $mc;
file_put_contents($output, $mc."n", FILE_APPEND);
echo 'url: ',$mc,' find ajax require so save it',PHP_EOL;
}
}
}
}
//check for javascript ajax require
echo 'remain: ',$num,' now start to check for javascript ajax require',PHP_EOL,PHP_EOL;
preg_match_all('//i', $result, $match);
if (CHECK_JS_TAG && isset($match[2]) && !empty($match[2]))
{
foreach ($match[2] as $mc)
{
$mc = trim($mc);
if ($mc == '#')
continue;
if (stripos($mc, 'http') === false)
$mc = $host.$mc;
if (($ret = check_valid_url($mc)))
{
//check for current page
foreach ($checkArr as $ck)
{
if (!isset($outputArr[md5($mc)]) && strpos($ret, $ck) !== false)
{
$outputArr[md5($mc)] = $mc;
file_put_contents($output, $mc."n", FILE_APPEND);
echo 'url: ',$mc,' find ajax require so save it',PHP_EOL;
break;
}
}
}
}
}
if ($reverse)
return;
//check for next page
preg_match_all('//i', $result, $match);
if (isset($match[2]) && !empty($match[2]))
{
echo 'check for next page, remain page counts: ',$num,PHP_EOL;
foreach ($match[2] as $mc)
{
$mc = trim($mc);
if ($mc == '#')
continue;
if (stripos($mc, 'http') === false)
$mc = $host.$mc;
echo 'check for next page: ',$mc,PHP_EOL;
work(check_valid_url($mc), true);
}
}
$num--;
sleep(3);
}
}
function check_valid_url($url)
{
if (stripos($url, 'http') === false)
$url = 'http://'.$url;
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, $url);
curl_setopt($ch, CURLOPT_HEADER, true);
curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)');
if (!is_null($header))
curl_setopt($ch, CURLOPT_HTTPHEADER, $header);
if (!is_null($proxy))
curl_setopt($ch, CURLOPT_PROXY, $proxy);
$ret = curl_exec($ch);
$errinfo = curl_error($ch);
curl_close($ch);
unset($ch);
if (!empty($errinfo) || ((strpos($ret, '200 OK') === false) && (strpos($ret, '302 Moved') === false)) || strpos($ret, '114so.cn') !== false)
{
showerror('check url: '.$url. ' find some errors', array($errinfo, $ret));
if (SAVE_ERROR)
file_put_contents(dirname(__FILE__).'/error.txt', $url."n", FILE_APPEND);
return false;
}
return $ret;
}
function check_extension()
{
if (!function_exists('curl_init') || !extension_loaded('curl'))
return false;
return true;
}
function showerror($t, $c)
{
$str = "#########################################################################n";
$str .= "# ".$t."n";
if (is_string($c))
$str .= "# ".$c;
elseif (is_array($c) && !empty($c))
{
foreach ($c as $c1)
$str .= "# ".$c1."n";
}
$str .= "n#########################################################################n";
echo $str;
unset($str);
}

复制代码

Statement

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

How do you modify data stored in a PHP session?Apr 27, 2025 am 12:23 AM

TomodifydatainaPHPsession,startthesessionwithsession_start(),thenuse$_SESSIONtoset,modify,orremovevariables.1)Startthesession.2)Setormodifysessionvariablesusing$_SESSION.3)Removevariableswithunset().4)Clearallvariableswithsession_unset().5)Destroythe

Give an example of storing an array in a PHP session.Apr 27, 2025 am 12:20 AM

Arrays can be stored in PHP sessions. 1. Start the session and use session_start(). 2. Create an array and store it in $_SESSION. 3. Retrieve the array through $_SESSION. 4. Optimize session data to improve performance.

How does garbage collection work for PHP sessions?Apr 27, 2025 am 12:19 AM

PHP session garbage collection is triggered through a probability mechanism to clean up expired session data. 1) Set the trigger probability and session life cycle in the configuration file; 2) You can use cron tasks to optimize high-load applications; 3) You need to balance the garbage collection frequency and performance to avoid data loss.

How can you trace session activity in PHP?Apr 27, 2025 am 12:10 AM

Tracking user session activities in PHP is implemented through session management. 1) Use session_start() to start the session. 2) Store and access data through the $_SESSION array. 3) Call session_destroy() to end the session. Session tracking is used for user behavior analysis, security monitoring, and performance optimization.

How can you use a database to store PHP session data?Apr 27, 2025 am 12:02 AM

Using databases to store PHP session data can improve performance and scalability. 1) Configure MySQL to store session data: Set up the session processor in php.ini or PHP code. 2) Implement custom session processor: define open, close, read, write and other functions to interact with the database. 3) Optimization and best practices: Use indexing, caching, data compression and distributed storage to improve performance.

Explain the concept of a PHP session in simple terms.Apr 26, 2025 am 12:09 AM

PHPsessionstrackuserdataacrossmultiplepagerequestsusingauniqueIDstoredinacookie.Here'showtomanagethemeffectively:1)Startasessionwithsession_start()andstoredatain$_SESSION.2)RegeneratethesessionIDafterloginwithsession_regenerate_id(true)topreventsessi

How do you loop through all the values stored in a PHP session?Apr 26, 2025 am 12:06 AM

In PHP, iterating through session data can be achieved through the following steps: 1. Start the session using session_start(). 2. Iterate through foreach loop through all key-value pairs in the $_SESSION array. 3. When processing complex data structures, use is_array() or is_object() functions and use print_r() to output detailed information. 4. When optimizing traversal, paging can be used to avoid processing large amounts of data at one time. This will help you manage and use PHP session data more efficiently in your actual project.

Explain how to use sessions for user authentication.Apr 26, 2025 am 12:04 AM

The session realizes user authentication through the server-side state management mechanism. 1) Session creation and generation of unique IDs, 2) IDs are passed through cookies, 3) Server stores and accesses session data through IDs, 4) User authentication and status management are realized, improving application security and user experience.

See all articles