How to use Python to implement job analysis reports-Python Tutorial-php.cn

Home

Backend Development

Python Tutorial

How to use Python to implement job analysis reports

WBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWBOYWB

May 01, 2023 pm 10:07 PM

python

1. The goal of this article

Get the Ajax request and parse the required fields in JSON

Save the data to Excel

Save the data to MySQL for easy analysis

2. Analysis results

1. Introduction of the library

Average salary levels of Python positions in five cities

2. Page structure

We enter the query The condition is Python as an example. Other conditions are not selected by default. Click Query to see all Python positions. Then we open the console and click the Network tab to see the following request:

How to use Python to implement job analysis reports

Judging from the response results, this request is exactly what we need. We can just request this address directly later. As can be seen from the picture, the following result is the information of each position.

Here we know where to request data and where to get the results. But there are only 15 pieces of data on the first page in the result list. How to get the data on other pages?

3. Request parameters

We click on the parameters tab, as follows:

We found that three form data were submitted. It is obvious that kd is the keyword we searched for. pn is the current page number. Just default to first, don't worry about it. All that's left is to construct a request to download 30 pages of data.

4. Constructing requests and parsing data

Constructing requests is very simple, we still use the requests library to do it. First, we construct the form data

data = {&#39;first&#39;: &#39;true&#39;, &#39;pn&#39;: page, &#39;kd&#39;: lang_name}

and then use requests to request the url address. The parsed JSON data is done. Since Lagou has strict restrictions on crawlers, we need to add all the headers fields in the browser and increase the crawler interval. I set it to 10-20s later, and then the data can be obtained normally.

import requests

def get_json(url, page, lang_name):
   headers = {
       &#39;Host&#39;: &#39;www.lagou.com&#39;,
       &#39;Connection&#39;: &#39;keep-alive&#39;,
       &#39;Content-Length&#39;: &#39;23&#39;,
       &#39;Origin&#39;: &#39;https://www.lagou.com&#39;,
       &#39;X-Anit-Forge-Code&#39;: &#39;0&#39;,
       &#39;User-Agent&#39;: &#39;Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:61.0) Gecko/20100101 Firefox/61.0&#39;,
       &#39;Content-Type&#39;: &#39;application/x-www-form-urlencoded; charset=UTF-8&#39;,
       &#39;Accept&#39;: &#39;application/json, text/javascript, */*; q=0.01&#39;,
       &#39;X-Requested-With&#39;: &#39;XMLHttpRequest&#39;,
       &#39;X-Anit-Forge-Token&#39;: &#39;None&#39;,
       &#39;Referer&#39;: &#39;https://www.lagou.com/jobs/list_python?city=%E5%85%A8%E5%9B%BD&cl=false&fromSearch=true&labelWords=&suginput=&#39;,
       &#39;Accept-Encoding&#39;: &#39;gzip, deflate, br&#39;,
       &#39;Accept-Language&#39;: &#39;en-US,en;q=0.9,zh-CN;q=0.8,zh;q=0.7&#39;
   }
   data = {&#39;first&#39;: &#39;false&#39;, &#39;pn&#39;: page, &#39;kd&#39;: lang_name}
   json = requests.post(url, data, headers=headers).json()
   list_con = json[&#39;content&#39;][&#39;positionResult&#39;][&#39;result&#39;]
   info_list = []
   for i in list_con:
       info = []
       info.append(i.get(&#39;companyShortName&#39;, &#39;无&#39;))
       info.append(i.get(&#39;companyFullName&#39;, &#39;无&#39;))
       info.append(i.get(&#39;industryField&#39;, &#39;无&#39;))
       info.append(i.get(&#39;companySize&#39;, &#39;无&#39;))
       info.append(i.get(&#39;salary&#39;, &#39;无&#39;))
       info.append(i.get(&#39;city&#39;, &#39;无&#39;))
       info.append(i.get(&#39;education&#39;, &#39;无&#39;))
       info_list.append(info)
   return info_list

4. Get all data

Now that we understand how to parse the data, the only thing left is to request all pages continuously. We construct a function to request all 30 pages of data.

def main():
   lang_name = &#39;python&#39;
   wb = Workbook()
   conn = get_conn()
   for i in [&#39;北京&#39;, &#39;上海&#39;, &#39;广州&#39;, &#39;深圳&#39;, &#39;杭州&#39;]:
       page = 1
       ws1 = wb.active
       ws1.title = lang_name
       url = &#39;https://www.lagou.com/jobs/positionAjax.json?city={}&needAddtionalResult=false&#39;.format(i)
       while page < 31:
           info = get_json(url, page, lang_name)
           page += 1
           import time
           a = random.randint(10, 20)
           time.sleep(a)
           for row in info:
               insert(conn, tuple(row))
               ws1.append(row)
   conn.close()
   wb.save(&#39;{}职位信息.xlsx&#39;.format(lang_name))

if __name__ == &#39;__main__&#39;:
   main()

The above is the detailed content of How to use Python to implement job analysis reports. For more information, please follow other related articles on the PHP Chinese website!

Statement

This article is reproduced at:亿速云. If there is any infringement, please contact admin@php.cn delete

Is Tuple Comprehension possible in Python? If yes, how and if not why?Apr 28, 2025 pm 04:34 PM

Article discusses impossibility of tuple comprehension in Python due to syntax ambiguity. Alternatives like using tuple() with generator expressions are suggested for creating tuples efficiently.(159 characters)

What are Modules and Packages in Python?Apr 28, 2025 pm 04:33 PM

The article explains modules and packages in Python, their differences, and usage. Modules are single files, while packages are directories with an __init__.py file, organizing related modules hierarchically.

What is docstring in Python?Apr 28, 2025 pm 04:30 PM

Article discusses docstrings in Python, their usage, and benefits. Main issue: importance of docstrings for code documentation and accessibility.

What is a lambda function?Apr 28, 2025 pm 04:28 PM

Article discusses lambda functions, their differences from regular functions, and their utility in programming scenarios. Not all languages support them.

What is a break, continue and pass in Python?Apr 28, 2025 pm 04:26 PM

Article discusses break, continue, and pass in Python, explaining their roles in controlling loop execution and program flow.

What is a pass in Python?Apr 28, 2025 pm 04:25 PM

The article discusses the 'pass' statement in Python, a null operation used as a placeholder in code structures like functions and classes, allowing for future implementation without syntax errors.

Can we Pass a function as an argument in Python?Apr 28, 2025 pm 04:23 PM

Article discusses passing functions as arguments in Python, highlighting benefits like modularity and use cases such as sorting and decorators.

What is the difference between / and // in Python?Apr 28, 2025 pm 04:21 PM

Article discusses / and // operators in Python: / for true division, // for floor division. Main issue is understanding their differences and use cases.Character count: 158

See all articles

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

Video Face Swap

Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

What's New in Windows 11 KB5054979 & How to Fix Update Issues

3 weeks agoByDDD

How to fix KB5055523 fails to install in Windows 11?

2 weeks agoByDDD

InZoi: How To Apply To School And University

3 weeks agoByDDD

How to fix KB5055518 fails to install in Windows 10?

2 weeks agoByDDD

Roblox: Dead Rails – How To Summon And Defeat Nikola Tesla

4 weeks agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

MantisBT

Mantis is an easy-to-deploy web-based defect tracking tool designed to aid in product defect tracking. It requires PHP, MySQL and a web server. Check out our demo and hosting services.

EditPlus Chinese cracked version

Small size, syntax highlighting, does not support code prompt function

SublimeText3 Chinese version

Chinese version, very easy to use

ZendStudio 13.5.1 Mac

Powerful PHP integrated development environment

SecLists

SecLists is the ultimate security tester's companion. It is a collection of various types of lists that are frequently used during security assessments, all in one place. SecLists helps make security testing more efficient and productive by conveniently providing all the lists a security tester might need. List types include usernames, passwords, URLs, fuzzing payloads, sensitive data patterns, web shells, and more. The tester can simply pull this repository onto a new test machine and he will have access to every type of list he needs.

Hot Topics

Where is the login entrance for gmail email?

7801

1644

1402

1299

1236