ITPub博客

首页 > 大数据 > Hadoop > hive中的Order By

hive中的Order By

Hadoop 作者:thamsyangsw 时间:2014-03-27 16:17:05 0 删除 编辑

hive中的order by也是对一个结果集合进行排序,但是和关系型数据库又所有不同。
这不同的地方也是两者在底层架构区别的体现。

hive的参数hive.mapred.mode是控制hive执行mapred的方式的,有两个选项:strict和nonstrict,默认值是nonstrict。
这个两个值对order by的执行有着很大的影响。

测试用例
hive> select * from test09;
OK
100     tom
200     mary
300     kate
400     tim
Time taken: 0.061 seconds

我们先来看看nonstrict的情况。

hive> set hive.mapred.mode=nonstrict;
hive> select * from test09 order by id;
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks determined at compile time: 1
In order to change the average load for a reducer (in bytes):
  set hive.exec.reducers.bytes.per.reducer=
In order to limit the maximum number of reducers:
  set hive.exec.reducers.max=
In order to set a constant number of reducers:
  set mapred.reduce.tasks=
Starting Job = job_201105020924_0065, Tracking URL = http://hadoop00:50030/jobdetails.jsp?jobid=job_201105020924_0065
Kill Command = /home/hjl/hadoop/bin/../bin/hadoop job  -Dmapred.job.tracker=hadoop00:9001 -kill job_201105020924_0065
2011-05-03 03:37:41,270 Stage-1 map = 0%,  reduce = 0%
2011-05-03 03:37:43,292 Stage-1 map = 50%,  reduce = 0%
2011-05-03 03:37:45,314 Stage-1 map = 100%,  reduce = 0%
2011-05-03 03:37:50,360 Stage-1 map = 100%,  reduce = 100%
Ended Job = job_201105020924_0065
OK
100     tom
200     mary
300     kate
400     tim
Time taken: 15.049 seconds

这个时候order by可以正常的执行,hive启动了一个reduce进行处理,事实上也只能启动一个reduce。

在来看看strict的情况
hive> set hive.mapred.mode=strict;
hive> select * from test09 order by id;
FAILED: Error in semantic analysis: line 1:30 In strict mode, limit must be specified if ORDER BY is present id

这个时候提示你,在strict模式下如果执行order by的操作必须要指定limit。
因为执行order by的时候只能启动单个reduce执行,如果排序的结果集过大,那么执行时间会非常漫长。

hive> select * from test09 order by id limit 4;
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks determined at compile time: 1
In order to change the average load for a reducer (in bytes):
  set hive.exec.reducers.bytes.per.reducer=
In order to limit the maximum number of reducers:
  set hive.exec.reducers.max=
In order to set a constant number of reducers:
  set mapred.reduce.tasks=
Starting Job = job_201105020924_0067, Tracking URL = http://hadoop00:50030/jobdetails.jsp?jobid=job_201105020924_0067
Kill Command = /home/hjl/hadoop/bin/../bin/hadoop job  -Dmapred.job.tracker=hadoop00:9001 -kill job_201105020924_0067
2011-05-03 04:18:26,828 Stage-1 map = 0%,  reduce = 0%
2011-05-03 04:18:27,842 Stage-1 map = 50%,  reduce = 0%
2011-05-03 04:18:29,864 Stage-1 map = 100%,  reduce = 0%
2011-05-03 04:18:35,916 Stage-1 map = 100%,  reduce = 100%
Ended Job = job_201105020924_0067
OK
100     tom
200     mary
300     kate
400     tim
Time taken: 15.706 seconds

加上limit后,SQL成功执行。

 

本文转自http://www.oratea.net/?p=622

来自 “ ITPUB博客 ” ,链接:http://blog.itpub.net/26613085/viewspace-1130847/,如需转载,请注明出处,否则将追究法律责任。

上一篇: hive中的sort by
下一篇: hive中的null值
请登录后发表评论 登录
全部评论

注册时间:2012-01-12

  • 博文量
    160
  • 访问量
    1176926