


How to get the correct number of applicants and viewers when crawling the 58.com work page?
58.com recruitment information crawling: Solve the problem of inconsistent data of applicants and viewers
When crawling the 58.com recruitment page, you often encounter a difficult problem: the number of applicants and the number of viewers displayed by the web page source code does not match the data actually displayed on the page, and the source code is often displayed as 0, while the data updated in real time on the page is consistent with the Elements content in the browser developer tool (F12). This article will explore how to solve this problem and obtain accurate applicants and viewers.
Problem analysis:
In order to prevent data from being maliciously crawled, 58.com adopted the method of dynamically loading data. The number of applicants and viewers on the page is not directly obtained from the HTML source code, but is loaded asynchronously through JavaScript. Therefore, direct parsing HTML source code cannot obtain the correct data.
Solution:
To obtain the correct number of applicants and viewers, you need to find the API interface provided by 58.com. By analyzing network requests, we can find an API interface for obtaining recruitment information statistics, with a URL similar to the following format:
<code>https://statisticszp.58.com/position/totalcount/?infoId=27988...</code>
The infoId
parameter represents the specific position ID and needs to be extracted based on the URL of the target recruitment page.
API returns data example:
The JSON data returned by the API interface contains the information we need:
{ "deliveryCount": 1141, // Number of applicants "commentCount": 0, "infoCount": 4, // Number of viewers "resumeReadPercent": 0, "referUrl": "", "nextUrl": "null" }
The deliveryCount
field indicates the number of applicants, and the infoCount
field indicates the number of viewers.
Implementation steps:
Get Job ID (infoId): Analyze the URL of the target recruitment page and find the parameter value corresponding to the Job ID. This may require the use of regular expressions or other string processing methods.
Construct API request URL: Replace the extracted
infoId
into the API URL template to form a complete API request URL.Send API requests: Use Python's
requests
library or other HTTP clients to send GET requests to the API URL.Analyze JSON data: parse the JSON data returned by the API into a Python dictionary, extract the values of
deliveryCount
andinfoCount
, that is, the correct number of applicants and number of viewers.
Through the above steps, you can bypass the dynamic loading mechanism of 58.com's web page and accurately obtain the number of applicants and viewers on the recruitment page. Please note that the address and parameter names of the API interface may change and need to be adjusted according to actual conditions. At the same time, please abide by 58.com's robots.txt rules to avoid excessive pressure on the server.
The above is the detailed content of How to get the correct number of applicants and viewers when crawling the 58.com work page?. For more information, please follow other related articles on the PHP Chinese website!

The future of HTML will develop in a more semantic, functional and modular direction. 1) Semanticization will make the tag describe the content more clearly, improving SEO and barrier-free access. 2) Functionalization will introduce new elements and attributes to meet user needs. 3) Modularity will support component development and improve code reusability.

HTMLattributesarecrucialinwebdevelopmentforcontrollingbehavior,appearance,andfunctionality.Theyenhanceinteractivity,accessibility,andSEO.Forexample,thesrcattributeintagsimpactsSEO,whileonclickintagsaddsinteractivity.Touseattributeseffectively:1)Usese

The alt attribute is an important part of the tag in HTML and is used to provide alternative text for images. 1. When the image cannot be loaded, the text in the alt attribute will be displayed to improve the user experience. 2. Screen readers use the alt attribute to help visually impaired users understand the content of the picture. 3. Search engines index text in the alt attribute to improve the SEO ranking of web pages.

The roles of HTML, CSS and JavaScript in web development are: 1. HTML is used to build web page structure; 2. CSS is used to beautify the appearance of web pages; 3. JavaScript is used to achieve dynamic interaction. Through tags, styles and scripts, these three together build the core functions of modern web pages.

Setting the lang attributes of a tag is a key step in optimizing web accessibility and SEO. 1) Set the lang attribute in the tag, such as. 2) In multilingual content, set lang attributes for different language parts, such as. 3) Use language codes that comply with ISO639-1 standards, such as "en", "fr", "zh", etc. Correctly setting the lang attribute can improve the accessibility of web pages and search engine rankings.

HTMLattributesareessentialforenhancingwebelements'functionalityandappearance.Theyaddinformationtodefinebehavior,appearance,andinteraction,makingwebsitesinteractive,responsive,andvisuallyappealing.Attributeslikesrc,href,class,type,anddisabledtransform

TocreatealistinHTML,useforunorderedlistsandfororderedlists:1)Forunorderedlists,wrapitemsinanduseforeachitem,renderingasabulletedlist.2)Fororderedlists,useandfornumberedlists,customizablewiththetypeattributefordifferentnumberingstyles.

HTML is used to build websites with clear structure. 1) Use tags such as, and define the website structure. 2) Examples show the structure of blogs and e-commerce websites. 3) Avoid common mistakes such as incorrect label nesting. 4) Optimize performance by reducing HTTP requests and using semantic tags.


Hot AI Tools

Undresser.AI Undress
AI-powered app for creating realistic nude photos

AI Clothes Remover
Online AI tool for removing clothes from photos.

Undress AI Tool
Undress images for free

Clothoff.io
AI clothes remover

Video Face Swap
Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

Hot Tools

PhpStorm Mac version
The latest (2018.2.1) professional PHP integrated development tool

DVWA
Damn Vulnerable Web App (DVWA) is a PHP/MySQL web application that is very vulnerable. Its main goals are to be an aid for security professionals to test their skills and tools in a legal environment, to help web developers better understand the process of securing web applications, and to help teachers/students teach/learn in a classroom environment Web application security. The goal of DVWA is to practice some of the most common web vulnerabilities through a simple and straightforward interface, with varying degrees of difficulty. Please note that this software

SublimeText3 Chinese version
Chinese version, very easy to use

SecLists
SecLists is the ultimate security tester's companion. It is a collection of various types of lists that are frequently used during security assessments, all in one place. SecLists helps make security testing more efficient and productive by conveniently providing all the lists a security tester might need. List types include usernames, passwords, URLs, fuzzing payloads, sensitive data patterns, web shells, and more. The tester can simply pull this repository onto a new test machine and he will have access to every type of list he needs.

Dreamweaver Mac version
Visual web development tools
