云数据库 SelectDB 版旨在提供卓越的性能和便捷的数据分析服务,在宽表聚合、多表关联以及高并发点查等场景下均具有优异的性能表现。本文将详细介绍SelectDB在TPC-H标准测试上的测试方法和测试结果。
概述
TPC-H是一个决策支持基准(Decision Support Benchmark),它由一套面向业务的特别查询和并发数据修改组成,使用的数据具有广泛的行业相关性。该基准测试通过一系列的查询操作来评估数据库系统在处理复杂查询和数据挖掘任务时的性能。TPC-H报告的性能指标称为TPC-H每小时复合查询性能指标(QphH@Size),该指标反映了系统处理查询能力的多个方面。这些方面包括执行查询时所选择的数据库大小,由单个流提交查询时的查询处理能力,以及由多个并发用户提交查询时的查询吞吐量等。
本文的TPC-H的实现基于TPC-H的基准测试,并不符合TPC-H基准测试的所有要求。本测试结果不能等同于完全遵守TPC-H测试规范所获得的测试结果,因此不能与完全遵守该测试规范获得的测试结果进行对比。
TPC-H 在内的标准测试集通常和实际业务场景差距较大,并且部分测试会针对测试集进行参数调优。所以标准测试集的测试结果仅能反映数据库在特定场景下的性能表现。建议使用实际业务数据进行进一步的测试。
测试环境
数据库环境。
环境配置项
配置说明
地域和可用区
华东1(杭州)地域,可用区K。
规格
64核 512 GB
磁盘
800 GB高性能云硬盘
云数据库 SelectDB 版内核版本
3.0.6
客户端环境。
环境配置项
配置说明
下载测试工具的设备
云服务器ECS实例,详情请参见创建实例。
地域和可用区
华东1(杭州)地域
实例规格
ecs.g7.2xlarge
操作系统
Ubuntu 22.04.1 LTS
网络
与云数据库 SelectDB 版实例为相同专有网络(VPC)。
测试数据集
整个测试模拟生成TPC-H 100GB和500GB的数据并导入SelectDB进行测试,下面是测试100GB数据表的相关说明及数据量。
TPC-H表名 | 行数 | 导入后大小 | 备注 |
REGION | 5 | 400KB | 区域表 |
NATION | 25 | 7.714 KB | 国家表 |
SUPPLIER | 100万 | 85.528 MB | 供应商表 |
PART | 2000万 | 752.330 MB | 零部件表 |
PARTSUPP | 8000万 | 4.375 GB | 零部件供应表 |
CUSTOMER | 1500万 | 1.317 GB | 客户表 |
ORDERS | 1.5亿 | 6.301 GB | 订单表 |
LINEITEM | 6亿 | 20.882 GB | 订单明细表 |
测试步骤
以下为测试的详细步骤。如何获取测试中涉及的脚本,请下载瑶池测试工具。
步骤一:安装unzip工具
安装unzip。
sudo apt install unzip
安装完成后,可以通过运行以下命令来验证unzip是否已成功安装。
unzip --help
步骤二:下载安装TPC-H数据生成工具
从上述脚本库中获取脚本后,解压脚本文件并进入对应目录,执行以下指令,下载并编译tpch-dbgen工具,示例如下。
tar -zxvf yaochi_performance_tool.tar.gz
cd ./yaochi_performance_tool/tpch-tools/bin
bash build-tpch-dbgen.sh
安装成功后,将在TPC-H_Tools_v3.0.0/
目录下生成dbgen
二进制文件。
步骤三:生成TPC-H测试集
在安装测试工具目录执行以下脚本生成TPC-H数据集,示例如下。
cd ./yaochi_performance_tool/tpch-tools/bin
bash gen-tpch-data.sh
数据会以.tbl
为后缀在tpch-data/
目录下生成,默认情况下的文件总大小约100GB。生成时间可能在数分钟到1小时不等。
步骤四:建表
准备
doris-cluster.conf
文件。在调用导入脚本前,需要将测试数据库的连接信息写在
doris-cluster.conf
文件中。文件位置在tpch-tools/conf/
目录下,内容包括连接集群的地址、HTTP端口、用户名、密码和待导入数据的DB。说明您可以在云数据库 SelectDB 版控制台的实例详情>网络信息中获取VPC地址(或公网地址)和HTTP协议端口。
export FE_HOST="xxx" export FE_HTTP_PORT="8080" export FE_QUERY_PORT="9030" export USER="root" export PASSWORD='xxx' export DB="tpch1"
创建TPC-H表。
在
./yaochi_performance_tool/tpch-tools/bin
目录下执行脚本来自动创建测试用表。bash create-tpch-tables.sh
步骤五:导入数据
在./yaochi_performance_tool/tpch-tools/bin
目录下执行脚本完成测试集数据导入。
bash ./load-tpch-data.sh
步骤六:检查导入数据
按照上述流程和参数执行的场合,数据量应和上文测试数据集给出的生成数据的行数一致。
SELECT COUNT(*) FROM lineitem;
SELECT COUNT(*) FROM orders;
SELECT COUNT(*) FROM partsupp;
SELECT COUNT(*) FROM part;
SELECT COUNT(*) FROM customer;
SELECT COUNT(*) FROM supplier;
SELECT COUNT(*) FROM nation;
SELECT COUNT(*) FROM region;
SELECT COUNT(*) FROM revenue0;
步骤七:查询测试
执行查询脚本
执行下面的命令完成查询测试。
./run-tpch-queries.sh
测试SQL详情请参见TPCH-Query-SQL。
说明目前SelectDB的查询优化器和统计信息功能仍有提升空间,所以我们在TPC-H中重写了一些查询以适应SelectDB的执行框架,但不影响结果的正确性。
单个 SQL 执行
下面是本次测试时使用的SQL语句,您也可以从代码库中获取最新的查询语句,最新测试查询语句请参见TPC-H 测试查询语句。
--ENV config SET GLOBAL experimental_enable_nereids_planner=true; SET GLOBAL experimental_enable_pipeline_engine=true; SET GLOBAL enable_runtime_filter_prune=false; SET GLOBAL runtime_filter_wait_time_ms=10000; SET GLOBAL enable_fallback_to_original_planner=false; SET GLOBAL query_timeout=1000; --Q1 SELECT l_returnflag, l_linestatus, sum(l_quantity) AS sum_qty, sum(l_extendedprice) AS sum_base_price, sum(l_extendedprice * (1 - l_discount)) AS sum_disc_price, sum(l_extendedprice * (1 - l_discount) * (1 + l_tax)) AS sum_charge, avg(l_quantity) AS avg_qty, avg(l_extendedprice) AS avg_price, avg(l_discount) AS avg_disc, count(*) AS count_order FROM lineitem WHERE l_shipdate <= date '1998-12-01' - interval '90' day GROUP BY l_returnflag, l_linestatus ORDER BY l_returnflag, l_linestatus; --Q2 SELECT s_acctbal, s_name, n_name, p_partkey, p_mfgr, s_address, s_phone, s_comment FROM partsupp join ( SELECT ps_partkey AS a_partkey, min(ps_supplycost) AS a_min FROM partsupp, part, supplier, nation, region WHERE p_partkey = ps_partkey AND s_suppkey = ps_suppkey AND s_nationkey = n_nationkey AND n_regionkey = r_regionkey AND r_name = 'EUROPE' AND p_size = 15 AND p_type LIKE '%BRASS' GROUP BY a_partkey ) AONps_partkey = a_partkey AND ps_supplycost=a_min , part, supplier, nation, region WHERE p_partkey = ps_partkey AND s_suppkey = ps_suppkey AND p_size = 15 AND p_type like '%BRASS' AND s_nationkey = n_nationkey AND n_regionkey = r_regionkey AND r_name = 'EUROPE' ORDER BY s_acctbal DESC, n_name, s_name, p_partkey limit 100; --Q3 SELECT l_orderkey, sum(l_extendedprice * (1 - l_discount)) AS revenue, o_orderdate, o_shippriority FROM ( SELECT l_orderkey, l_extendedprice, l_discount, o_orderdate, o_shippriority, o_custkey FROM lineitem JOIN orders WHERE l_orderkey = o_orderkey AND o_orderdate < date '1995-03-15' AND l_shipdate > date '1995-03-15' ) t1 JOIN customer c ONc.c_custkey = t1.o_custkey WHERE c_mktsegment = 'BUILDING' GROUP BY l_orderkey, o_orderdate, o_shippriority ORDER BY revenue DESC, o_orderdate limit 10; --Q4 SELECT o_orderpriority, count(*) AS order_count FROM ( SELECT * FROM lineitem WHERE l_commitdate < l_receiptdate ) t1 RIGHT semi JOIN orders ONt1.l_orderkey = o_orderkey WHERE o_orderdate >= date '1993-07-01' AND o_orderdate < date '1993-07-01' + interval '3' month GROUP BY o_orderpriority ORDER BY o_orderpriority; --Q5 SELECT n_name, sum(l_extendedprice * (1 - l_discount)) AS revenue FROM customer, orders, lineitem, supplier, nation, region WHERE c_custkey = o_custkey AND l_orderkey = o_orderkey AND l_suppkey = s_suppkey AND c_nationkey = s_nationkey AND s_nationkey = n_nationkey AND n_regionkey = r_regionkey AND r_name = 'ASIA' AND o_orderdate >= date '1994-01-01' AND o_orderdate < date '1994-01-01' + interval '1' year GROUP BY n_name ORDER BY revenue DESC; --Q6 SELECT sum(l_extendedprice * l_discount) AS revenue FROM lineitem WHERE l_shipdate >= date '1994-01-01' AND l_shipdate < date '1994-01-01' + interval '1' year AND l_discount between .06 - 0.01 AND .06 + 0.01 AND l_quantity < 24; --Q7 SELECT supp_nation, cust_nation, l_year, sum(volume) AS revenue FROM ( SELECT n1.n_name AS supp_nation, n2.n_name AS cust_nation, extract(year FROM l_shipdate) AS l_year, l_extendedprice * (1 - l_discount) AS volume FROM supplier, lineitem, orders, customer, nation n1, nation n2 WHERE s_suppkey = l_suppkey AND o_orderkey = l_orderkey AND c_custkey = o_custkey AND s_nationkey = n1.n_nationkey AND c_nationkey = n2.n_nationkey AND ( (n1.n_name = 'FRANCE' AND n2.n_name = 'GERMANY') or (n1.n_name = 'GERMANY' AND n2.n_name = 'FRANCE') ) AND l_shipdate between date '1995-01-01' AND date '1996-12-31' ) AS shipping GROUP BY supp_nation, cust_nation, l_year ORDER BY supp_nation, cust_nation, l_year; --Q8 SELECT o_year, sum(case when nation = 'BRAZIL' then volume else 0 end) / sum(volume) AS mkt_share FROM ( SELECT extract(year FROM o_orderdate) AS o_year, l_extendedprice * (1 - l_discount) AS volume, n2.n_name AS nation FROM lineitem, orders, customer, supplier, part, nation n1, nation n2, region WHERE p_partkey = l_partkey AND s_suppkey = l_suppkey AND l_orderkey = o_orderkey AND o_custkey = c_custkey AND c_nationkey = n1.n_nationkey AND n1.n_regionkey = r_regionkey AND r_name = 'AMERICA' AND s_nationkey = n2.n_nationkey AND o_orderdate between date '1995-01-01' AND date '1996-12-31' AND p_type = 'ECONOMY ANODIZED STEEL' ) AS all_nations GROUP BY o_year ORDER BY o_year; --Q9 SELECT nation, o_year, sum(amount) AS sum_profit FROM ( SELECT n_name AS nation, extract(year FROM o_orderdate) AS o_year, l_extendedprice * (1 - l_discount) - ps_supplycost * l_quantity AS amount FROM lineitem JOIN ordersONo_orderkey = l_orderkey join[shuffle] partONp_partkey = l_partkey join[shuffle] partsuppONps_partkey = l_partkey join[shuffle] supplierONs_suppkey = l_suppkey join[broadcast] nationONs_nationkey = n_nationkey WHERE ps_suppkey = l_suppkey AND p_name like '%green%' ) AS profit GROUP BY nation, o_year ORDER BY nation, o_year DESC; --Q10 SELECT c_custkey, c_name, sum(t1.l_extendedprice * (1 - t1.l_discount)) AS revenue, c_acctbal, n_name, c_address, c_phone, c_comment FROM customer, ( SELECT o_custkey,l_extendedprice,l_discount FROM lineitem, orders WHERE l_orderkey = o_orderkey AND o_orderdate >= date '1993-10-01' AND o_orderdate < date '1993-10-01' + interval '3' month AND l_returnflag = 'R' ) t1, nation WHERE c_custkey = t1.o_custkey AND c_nationkey = n_nationkey GROUP BY c_custkey, c_name, c_acctbal, c_phone, n_name, c_address, c_comment ORDER BY revenue DESC limit 20; --Q11 SELECT ps_partkey, sum(ps_supplycost * ps_availqty) AS value FROM partsupp, ( SELECT s_suppkey FROM supplier, nation WHERE s_nationkey = n_nationkey AND n_name = 'GERMANY' ) B WHERE ps_suppkey = B.s_suppkey GROUP BY ps_partkey having sum(ps_supplycost * ps_availqty) > ( SELECT sum(ps_supplycost * ps_availqty) * 0.000002 FROM partsupp, (SELECT s_suppkey FROM supplier, nation WHERE s_nationkey = n_nationkey AND n_name = 'GERMANY' ) A WHERE ps_suppkey = A.s_suppkey ) ORDER BY value DESC; --Q12 SELECT l_shipmode, sum(case when o_orderpriority = '1-URGENT' or o_orderpriority = '2-HIGH' then 1 else 0 end) AS high_line_count, sum(case when o_orderpriority <> '1-URGENT' AND o_orderpriority <> '2-HIGH' then 1 else 0 end) AS low_line_count FROM orders, lineitem WHERE o_orderkey = l_orderkey AND l_shipmode in ('MAIL', 'SHIP') AND l_commitdate < l_receiptdate AND l_shipdate < l_commitdate AND l_receiptdate >= date '1994-01-01' AND l_receiptdate < date '1994-01-01' + interval '1' year GROUP BY l_shipmode ORDER BY l_shipmode; --Q13 SELECT c_count, count(*) AS custdist FROM ( SELECT c_custkey, count(o_orderkey) AS c_count FROM orders RIGHT outer JOIN customer on c_custkey = o_custkey AND o_comment not like '%special%requests%' GROUP BY c_custkey ) AS c_orders GROUP BY c_count ORDER BY custdist DESC, c_count DESC; --Q14 SELECT 100.00 * sum(case when p_type like 'PROMO%' then l_extendedprice * (1 - l_discount) else 0 end) / sum(l_extendedprice * (1 - l_discount)) AS promo_revenue FROM part, lineitem WHERE l_partkey = p_partkey AND l_shipdate >= date '1995-09-01' AND l_shipdate < date '1995-09-01' + interval '1' month; --Q15 SELECT s_suppkey, s_name, s_address, s_phone, total_revenue FROM supplier, revenue0 WHERE s_suppkey = supplier_no AND total_revenue = ( SELECT max(total_revenue) FROM revenue0 ) ORDER BY s_suppkey; --Q16 SELECT p_brAND, p_type, p_size, count(distinct ps_suppkey) AS supplier_cnt FROM partsupp, part WHERE p_partkey = ps_partkey AND p_brAND <> 'BrAND#45' AND p_type not like 'MEDIUM POLISHED%' AND p_size in (49, 14, 23, 45, 19, 3, 36, 9) AND ps_suppkey not in ( SELECT s_suppkey FROM supplier WHERE s_comment like '%Customer%Complaints%' ) GROUP BY p_brAND, p_type, p_size ORDER BY supplier_cnt DESC, p_brAND, p_type, p_size; --Q17 SELECT sum(l_extendedprice) / 7.0 AS avg_yearly FROM lineitem JOIN [broadcast] part p1ONp1.p_partkey = l_partkey WHERE p1.p_brAND = 'BrAND#23' AND p1.p_container = 'MED BOX' AND l_quantity < ( SELECT 0.2 * avg(l_quantity) FROM lineitem JOIN [broadcast] part p2ONp2.p_partkey = l_partkey WHERE l_partkey = p1.p_partkey AND p2.p_brAND = 'BrAND#23' AND p2.p_container = 'MED BOX' ); --Q18 SELECT c_name, c_custkey, t3.o_orderkey, t3.o_orderdate, t3.o_totalprice, sum(t3.l_quantity) FROM customer join ( SELECT * FROM lineitem join ( SELECT * FROM orders left semi join ( SELECT l_orderkey FROM lineitem GROUP BY l_orderkey having sum(l_quantity) > 300 ) t1 ONo_orderkey = t1.l_orderkey ) t2 ONt2.o_orderkey = l_orderkey ) t3 on c_custkey = t3.o_custkey GROUP BY c_name, c_custkey, t3.o_orderkey, t3.o_orderdate, t3.o_totalprice ORDER BY t3.o_totalprice DESC, t3.o_orderdate limit 100; --Q19 SELECT sum(l_extendedprice* (1 - l_discount)) AS revenue FROM lineitem, part WHERE ( p_partkey = l_partkey AND p_brAND = 'BrAND#12' AND p_container in ('SM CASE', 'SM BOX', 'SM PACK', 'SM PKG') AND l_quantity >= 1 AND l_quantity <= 1 + 10 AND p_size between 1 AND 5 AND l_shipmode in ('AIR', 'AIR REG') AND l_shipinstruct = 'DELIVER IN PERSON' ) or ( p_partkey = l_partkey AND p_brAND = 'BrAND#23' AND p_container in ('MED BAG', 'MED BOX', 'MED PKG', 'MED PACK') AND l_quantity >= 10 AND l_quantity <= 10 + 10 AND p_size between 1 AND 10 AND l_shipmode in ('AIR', 'AIR REG') AND l_shipinstruct = 'DELIVER IN PERSON' ) or ( p_partkey = l_partkey AND p_brAND = 'BrAND#34' AND p_container in ('LG CASE', 'LG BOX', 'LG PACK', 'LG PKG') AND l_quantity >= 20 AND l_quantity <= 20 + 10 AND p_size between 1 AND 15 AND l_shipmode in ('AIR', 'AIR REG') AND l_shipinstruct = 'DELIVER IN PERSON' ); --Q20 SELECT s_name, s_address FROM supplier left semi join ( SELECT * FROM ( SELECT l_partkey,l_suppkey, 0.5 * sum(l_quantity) AS l_q FROM lineitem WHERE l_shipdate >= date '1994-01-01' AND l_shipdate < date '1994-01-01' + interval '1' year GROUP BY l_partkey,l_suppkey ) t2 join ( SELECT ps_partkey, ps_suppkey, ps_availqty FROM partsupp left semi JOIN part ONps_partkey = p_partkey AND p_name like 'forest%' ) t1 ONt2.l_partkey = t1.ps_partkey AND t2.l_suppkey = t1.ps_suppkey AND t1.ps_availqty > t2.l_q ) t3 on s_suppkey = t3.ps_suppkey join nation WHERE s_nationkey = n_nationkey AND n_name = 'CANADA' ORDER BY s_name; --Q21 SELECT s_name, count(*) AS numwait FROM lineitem l2 RIGHT semi join ( SELECT * FROM lineitem l3 RIGHT anti join ( SELECT * FROM orders JOIN lineitem l1ONl1.l_orderkey = o_orderkey AND o_orderstatus = 'F' join ( SELECT * FROM supplier JOIN nation WHERE s_nationkey = n_nationkey AND n_name = 'SAUDI ARABIA' ) t1 WHERE t1.s_suppkey = l1.l_suppkey AND l1.l_receiptdate > l1.l_commitdate ) t2 ONl3.l_orderkey = t2.l_orderkey AND l3.l_suppkey <> t2.l_suppkey AND l3.l_receiptdate > l3.l_commitdate ) t3 ONl2.l_orderkey = t3.l_orderkey AND l2.l_suppkey <> t3.l_suppkey GROUP BY t3.s_name ORDER BY numwait DESC, t3.s_name limit 100; --Q22 with tmp AS (SELECT avg(c_acctbal) AS av FROM customer WHERE c_acctbal > 0.00 AND substring(c_phone, 1, 2) in ('13', '31', '23', '29', '30', '18', '17')) SELECT cntrycode, count(*) AS numcust, sum(c_acctbal) AS totacctbal FROM ( SELECT substring(c_phone, 1, 2) AS cntrycode, c_acctbal FROM orders RIGHT anti JOIN customer cON o_custkey = c.c_custkey JOIN tmpONc.c_acctbal > tmp.av WHERE substring(c_phone, 1, 2) in ('13', '31', '23', '29', '30', '18', '17') ) AS custsale GROUP BY cntrycode ORDER BY cntrycode;
测试结果
以下为TPCH 100 GB和500 GB的测试结果。
Query | TPCH 100GB(s) | TPCH 500GB(s) |
Q1 | 1.74 | 10.04 |
Q2 | 0.07 | 0.19 |
Q3 | 0.34 | 3.43 |
Q4 | 0.19 | 1.1 |
Q5 | 0.81 | 7.52 |
Q6 | 0.03 | 0.15 |
Q7 | 0.54 | 5.74 |
Q8 | 0.26 | 2.56 |
Q9 | 2.62 | 18.44 |
Q10 | 0.91 | 5.45 |
Q11 | 0.08 | 0.36 |
Q12 | 0.09 | 0.47 |
Q13 | 1.32 | 6.6 |
Q14 | 0.12 | 0.59 |
Q15 | 0.18 | 0.85 |
Q16 | 0.28 | 1.17 |
Q17 | 0.1 | 0.45 |
Q18 | 1.7 | 9.94 |
Q19 | 0.18 | 1.9 |
Q20 | 0.39 | 0.62 |
Q21 | 0.65 | 7.23 |
Q22 | 0.19 | 1.04 |
合计 | 12.79 | 85.84 |