
Pandas 的 describe() 默认使用线性插值法(interpolation="linear")计算四分位数,其核心公式为 i + (j - i) * fraction,其中 i 和 j 是相邻两个数据点,fraction 是目标位置的小数部分。
pandas 的 `describe()` 默认使用线性插值法(`interpolation="linear"`)计算四分位数,其核心公式为 `i + (j - i) * fraction`,其中 `i` 和 `j` 是相邻两个数据点,`fraction` 是目标位置的小数部分。
在 Pandas 中,Series.describe() 或 DataFrame.describe() 输出的 25%、50%(中位数)、75% 等统计量,本质调用的是底层 quantile() 方法,默认采用 interpolation="linear" 插值策略。理解其计算逻辑对结果复现与数据解释至关重要。
以示例数据 [10, 13, 15, 19, 21, 25](共 n = 6 个有序观测值)为例,计算第 25 百分位数(即第一四分位数 Q1):
- Pandas 将百分位位置映射到排序后数组的索引区间:对于 q = 0.25,其理论位置为 pos = q × (n − 1) = 0.25 × 5 = 1.25;
- 该位置落在索引 1 和 2 之间(对应值 13 和 15),整数部分 i = 1 → 值 13,j = 2 → 值 15;
- 小数部分 fraction = 0.25;
- 代入线性插值公式:
13 + (15 − 13) × 0.25 = 13 + 2 × 0.25 = 13.5。
这正是 s.describe() 输出中 25% 为 13.500000 的由来。
你可通过显式调用 quantile() 验证:
import pandas as pd s = pd.Series([10, 13, 15, 19, 21, 25]) print(s.quantile(0.25, interpolation="linear")) # 输出: 13.5 print(s.quantile([0.25, 0.5, 0.75], interpolation="linear")) # 输出: # 0.25 13.5 # 0.50 17.0 # 0.75 20.5
⚠️ 注意事项:
- interpolation 参数在 Pandas 1.1.0+ 已弃用,推荐改用 method 参数(如 method="linear"),但 describe() 内部仍默认等价于旧版 "linear" 行为;
- 其他常用插值方式包括 "lower"(向下取)、"higher"(向上取)、"nearest"(最近邻)、"midpoint"(中点)等,结果差异显著;
- describe() 的百分位计算不等同于 Excel 的 PERCENTILE.INC(默认)或 PERCENTILE.EXC,也不完全匹配 R 或 NumPy 的默认策略——务必确认工具链一致性;
- 若需与传统教科书定义(如“将数据分为四等份”)对齐,建议明确指定 method="nearest" 或预处理数据,而非依赖默认行为。
总结:Pandas 的 describe() 百分位数是基于有序样本、按 q × (n−1) 定位、再线性插值得出的连续估计值,其设计兼顾平滑性与统计稳健性,但在学术报告或跨平台对比时,应主动声明所用插值方法以确保可复现性。










