全角字符的判断依据是unicode码点是否落在0xff01–0xff5e、0x3000、0x3001–0x303f、0x3040–0x309f、0x30a0–0x30ff、0x4e00–0x9fff等区间,需先将utf-8字符串解码为码点再判断。

全角字符的判断依据是什么
全角字符在 Unicode 中主要分布在 0xFF01–0xFF5E(全角 ASCII 可见字符)、0x3000(全角空格)、0x3001–0x303F(标点符号)、0x4E00–0x9FFF(中日韩汉字)等区间。C++ 标准库不提供“是否全角”的内置函数,必须手动检查码点。
注意:std::string 存储的是 UTF-8 编码字节流,直接遍历 char 会出错;必须先解码为 Unicode 码点(如用 std::wstring + std::locale,或更可靠地用第三方库如 ICU、utf8cpp,或手写 UTF-8 解码逻辑)。
用 utf8cpp 解码后逐码点判断
推荐轻量方案:使用 utf8cpp(header-only),它能安全将 std::string(UTF-8)转为 std::vector<uint32_t></uint32_t> 码点序列。
- 安装:仅需包含
utf8.h - 判断逻辑:对每个
uint32_t码点,检查是否落在典型全角区间内 - 常见误判点:全角数字
0(U+FF10)和全角字母A(U+FF21)属于0xFF01–0xFF5E;但全角平假名、片假名(如あU+3042)在0x3040–0x309F,也要纳入
示例代码片段:
#include "utf8.h"
#include <string>
#include <vector>
bool has_fullwidth(const std::string& s) {
std::vector<uint32_t> cp;
utf8::utf8to32(s.begin(), s.end(), back_inserter(cp));
for (uint32_t c : cp) {
if ((c >= 0xFF01 && c = 0x3001 && c = 0x3040 && c = 0x30A0 && c = 0x4E00 && c
<h3>不依赖外部库的手动 UTF-8 解码(慎用)</h3>
<p>若无法引入 utf8cpp,需自己解析 UTF-8 字节序列。但极易出错——比如把多字节序列当单字节处理,导致码点错位。</p><div class="aritcle_card flexRow artxards">
<div class="artcardd flexRow">
<a class="aritcle_card_img" rel="nofollow" href="/xiazai/skill5502" title="C++ Code Review Master"><img
src="https://img.php.cn/upload/skill/000/000/081/179051228971575.jpg" alt="C++ Code Review Master" onerror="this.onerror='';this.src='/static/lhimages/moren/morentu.png'" ></a>
<div class="aritcle_card_info flexColumn">
<a rel="nofollow" href="/xiazai/skill5502" title="C++ Code Review Master" class="overflowclass">C++ Code Review Master</a>
<p class="overflowclass">组合式C++代码评审方案,融合静态分析、AI推理、多轮迭代评审和C++专项检查,适用于PR审查、增量代码审查、全项目评审和代码质量评分,触发词包括review cpp、cpp代码评审、C++review、代码审查。</p>
</div>
<a rel="nofollow" href="/xiazai/skill5502" title="C++ Code Review Master" class="aritcle_card_btn flexRow flexcenter"><b></b><span>下载</span>
</a>
</div>
</div>
<p>关键点:</p>
<ul>
<li>UTF-8 中,首字节以 <code>0xxxxxxx</code> 开头是 ASCII(1 字节),<code>110xxxxx</code> 是 2 字节,<code>1110xxxx</code> 是 3 字节,<code>11110xxx</code> 是 4 字节</li>
<li>后续字节必须是 <code>10xxxxxx</code>,否则非法;跳过非法序列会导致后续全部错乱</li>
<li>
<code>std::string::iterator</code> 遍历 <code>char</code> 不能直接用于判断,必须按 UTF-8 规则步进</li>
</ul>
<p>简单但脆弱的实现示例(仅作理解,生产环境请用 utf8cpp):</p>
<pre class="brush:php;toolbar:false;">
bool is_fullwidth_codepoint(uint32_t cp) {
return (cp >= 0xFF01 && cp = 0x3001 && cp = 0x3040 && cp = 0x30A0 && cp = 0x4E00 && cp
<h3>为什么不能用 locale 或 wstring 直接判断</h3>
<p>很多人尝试用 <code>std::wstring_convert<:codecvt_utf8>></:codecvt_utf8></code> 转成 <code>wstring</code> 再遍历 <code>wchar_t</code>,但这是危险的:</p>
- Windows 下
wchar_t是 UTF-16,遇到代理对(surrogate pair)时单个wchar_t无法表示一个完整码点 - Linux/macOS 下
wchar_t通常是 UTF-32,看似安全,但std::codecvt_utf8<wchar_t></wchar_t>在 C++17 已被弃用,且不同标准库实现行为不一致 -
std::locale的ctype<char>::is(ctype_base::space, c)</char>等只识别基础分类,不区分全半角
真正可靠的路径只有一条:从 UTF-8 字节串出发,正确解码为 Unicode 码点,再查表判断。其余捷径几乎都埋着兼容性或正确性雷。
C++免费学习笔记(深入):立即使用
在学习笔记中,你将探索 C++ 的入门与实战技巧!










