Home  >  Article  >  Backend Development  >  Application examples of PHP CURL library_PHP tutorial

Application examples of PHP CURL library_PHP tutorial

WBOY
WBOYOriginal
2016-07-13 10:33:26816browse

Use PHP’s cURL library to crawl web pages simply and effectively. You only need to run a script and analyze the web pages you crawled, and then you can get the data you want programmatically. Whether you want to retrieve partial data from a link, take an XML file and import it into a database, or even simply retrieve the content of a web page, cURL is a powerful PHP library.

The commonly used functions of the CURL function library (Client URL Library Function) in PHP are as follows:

  • curl_close — close a curl session
  • curl_copy_handle — Copy all contents and parameters of a curl connection resource
  • curl_errno — Returns a numeric number containing error information for the current session
  • curl_error — Returns a string containing error information for the current session
  • curl_exec — execute a curl session
  • curl_getinfo — Get information about a curl connection resource handle
  • curl_init — initialize a curl session
  • curl_multi_add_handle — Add individual curl handle resources to a curl batch session
  • curl_multi_close — close a batch handle resource
  • curl_multi_exec — parse a curl batch handle
  • curl_multi_getcontent — Returns the text stream of the obtained output
  • curl_multi_info_read — Get the relevant transmission information of the currently parsed curl
  • curl_multi_init — Initialize a curl batch handle resource
  • curl_multi_remove_handle — remove a handle resource in the curl batch handle resource
  • curl_multi_select — Get all the sockets associated with the cURL extension, which can then be "selected"
  • curl_setopt_array — Set session parameters for a curl in the form of an array
  • curl_setopt — Set session parameters for a curl
  • curl_version — Get curl-related version information
  • The role of the curl_init() function is to initialize a curl session. The only parameter of the curl_init() function is optional and represents a URL address.
  • The curl_exec() function is used to execute a curl session. The only parameter is the handle returned by the curl_init() function.
  • The curl_close() function is used to close a curl session. The only parameter is the handle returned by the curl_init() function.

Basic example

<?php
// 初始化一个 cURL 对象
$curl = curl_init();

// 设置你需要抓取的URL
curl_setopt($curl, CURLOPT_URL, 'http://www.cmx8.cn');

// 设置header
curl_setopt($curl, CURLOPT_HEADER, 1);

// 设置cURL 参数,要求结果保存到字符串中还是输出到屏幕上。
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1);

// 运行cURL,请求网页
$data = curl_exec($curl);

// 关闭URL请求
curl_close($curl);

// 显示获得的数据
var_dump($data);
?>

POST data

Accept two form fields, one is the phone number and the other is the text message content.

<?php
$phoneNumber = '13812345678';
$message = 'This message was generated by curl and php';
$curlPost = 'pNUMBER=' . urlencode($phoneNumber) . '&MESSAGE=' . urlencode($message) . '&SUBMIT=Send';
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, 'http://www.lxvoip.com/sendSMS.php');
curl_setopt($ch, CURLOPT_HEADER, 1);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($ch, CURLOPT_POST, 1);
curl_setopt($ch, CURLOPT_POSTFIELDS, $curlPost);
$data = curl_exec();
curl_close($ch);
?>

Use a proxy server

<?php
$ch = curl_init();
curl_setopt($ch, CURLOPT_URL, 'http://www.cmx8.cn');
curl_setopt($ch, CURLOPT_HEADER, 1);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
curl_setopt($ch, CURLOPT_HTTPPROXYTUNNEL, 1);
curl_setopt($ch, CURLOPT_PROXY, 'proxy.lxvoip.com:1080');
curl_setopt($ch, CURLOPT_PROXYUSERPWD, 'user:password');
$data = curl_exec();
curl_close($ch);
?>

Simulated login

Simulate login to discuz program.

<?php
/**   
* Curl 模拟登录 discuz 程序   
* 尚未实现开启验证码的的论坛登录功能   
*/   
   
!extension_loaded('curl') && die('The curl extension is not loaded.');    
   
$discuz_url = 'http://www.lxvoip.com';//论坛地址    
$login_url = $discuz_url .'/logging.php?action=login';//登录页地址    
$get_url = $discuz_url .'/my.php?item=threads'; //我的帖子    
   
$post_fields = array();    
//以下两项不需要修改    
$post_fields['loginfield'] = 'username';    
$post_fields['loginsubmit'] = 'true';    
//用户名和密码,必须填写    
$post_fields['username'] = 'lxvoip';    
$post_fields['password'] = '88888888';    
//安全提问    
$post_fields['questionid'] = 0;    
$post_fields['answer'] = '';    
//@todo验证码    
$post_fields['seccodeverify'] = '';    
   
//获取表单FORMHASH    
$ch = curl_init($login_url);    
curl_setopt($ch, CURLOPT_HEADER, 0);    
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);    
$contents = curl_exec($ch);    
curl_close($ch);    
preg_match('/<input\s*type="hidden"\s*name="formhash"\s*value="(.*?)"\s*\ />/i', $contents, $matches);    
if(!empty($matches)) {    
    $formhash = $matches[1];    
} else {    
    die('Not found the forumhash.');    
}    
   
//POST数据,获取COOKIE    
$cookie_file = dirname(__FILE__) . '/cookie.txt';    
//$cookie_file = tempnam('/tmp');    
$ch = curl_init($login_url);    
curl_setopt($ch, CURLOPT_HEADER, 0);    
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);    
curl_setopt($ch, CURLOPT_POST, 1);    
curl_setopt($ch, CURLOPT_POSTFIELDS, $post_fields);    
curl_setopt($ch, CURLOPT_COOKIEJAR, $cookie_file);    
curl_exec($ch);    
curl_close($ch);    
   
//带着上面得到的COOKIE获取需要登录后才能查看的页面内容    
$ch = curl_init($get_url);    
curl_setopt($ch, CURLOPT_HEADER, 0);    
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 0);    
curl_setopt($ch, CURLOPT_COOKIEFILE, $cookie_file);    
$contents = curl_exec($ch);    
curl_close($ch);    
   
var_dump($contents); 
?>

www.bkjia.comtruehttp: //www.bkjia.com/PHPjc/752496.htmlTechArticleUsing PHP’s cURL library can easily and effectively capture web pages. You only need to run a script and analyze the web pages you crawled, and then you can get what you want programmatically...
Statement:
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn