html解析引擎边下载边解析,通过状态机将字符流转为token,再用栈机制构建dom树,具备高容错性。

HTML解析引擎不是等整个文件下载完才开始干活,而是边收字节边建节点树——这是理解整个过程的关键前提。
字符流 → Token 是状态机在“读”
浏览器拿到HTTP响应体后,第一件事是按编码(如UTF-8)把字节解成Unicode字符;接着用一个预定义的状态机逐个“吃”字符。遇到 就切到“标签开始”状态,后面跟字母就认定为<code>startTag,跟/就转为endTag,遇到空格或换行则可能进入“属性名”或“文本”状态。
常见错误现象:Uncaught DOMException: Failed to execute 'insertAdjacentHTML' on 'Element' 往往是因为传入了未闭合的片段(比如只写了 <div>),导致后续Token序列错乱,状态机卡住或误判。<p>这个阶段不关心语义是否合法,只保证:
- <code><img src="a.jpg"> 被拆成 startTag(name=“img”,attrs=[{name:“src”, value:“a.jpg”}])
- <!-- comment --> 被识别为 comment 类型Token,直接丢弃
- 连续空白符(含换行)在文本Token中被压缩为单个空格,除非父元素是 <pre class="brush:php;toolbar:false;"></pre> 或设了 white-space: pre
Token → 节点 → DOM树 是栈在“管”嵌套
每个startTag触发新建节点,并压入一个开放元素栈(open element stack);每个endTag则尝试弹出栈顶——但不是无条件弹,而是按HTML规范做容错匹配。例如 <div><p></p><div class="aritcle_card flexRow artxards">
<div class="artcardd flexRow">
<a class="aritcle_card_img" rel="nofollow" href="/xiazai/skill5806" title="html-deploy"><img
src="https://img.php.cn/upload/skill/000/000/081/179066538882434.jpg" alt="html-deploy" onerror="this.onerror='';this.src='/static/lhimages/moren/morentu.png'" ></a>
<div class="aritcle_card_info flexColumn">
<a rel="nofollow" href="/xiazai/skill5806" title="html-deploy" class="overflowclass">html-deploy</a>
<p class="overflowclass">使用 htmlcode.fun 将 HTML 内容或文件部署到网页,适用于用户要求“部署到网页”“托管此 HTML”“生成此前端...的实时链接”等场景。</p>
</div>
<a rel="nofollow" href="/xiazai/skill5806" title="html-deploy" class="aritcle_card_btn flexRow flexcenter"><b></b><span>下载</span>
</a>
</div>
</div></div> 中,解析器发现










