spark core sortBy和sortByKey探索

2023-07-22 04:57:09

感覺自己好久沒有更新過部落格了，本人最近有點兒迷失，特來寫篇技術部落格，以做自警

不知道大家有沒有注意到，大家在編寫spark程式調用sortBy/sortByKey這兩個算子的時候大家會不會有這樣子的疑問，他們兩個明明是transformation，為啥在執行的時候卻觸發了作業的執行呢？今天就和大家一起一探究竟？

val wordCountRdd = spark.sparkContext.textFile(path).
      flatMap(_.split(" ")).
      map(word => (word, 1)).
      reduceByKey(_ + _)

    val sortByCountDescRdd = wordCountRdd.sortBy(-_._2)

當你在shell輸入一下的code時，會發現如下圖：

spark core sortBy和sortByKey探索

出現了類似action的運作條，到底怎麼回事兒呢？

首先sortBy的實作就是調用了sortByKey,是以我們隻關注sortByKey的實作

def sortByKey(ascending: Boolean = true, numPartitions: Int = self.partitions.length)
      : RDD[(K, V)] = self.withScope
  {
    val part = new RangePartitioner(numPartitions, self, ascending)
    new ShuffledRDD[K, V, V](self, part)
      .setKeyOrdering(if (ascending) ordering else ordering.reverse)
  }

此時先看RangePartitioner分區器的實作，隻挑重要部分開始描述啊

private[spark] object RangePartitioner {

  /**
   * Sketches the input RDD via reservoir sampling on each partition.
   *
   * @param rdd the input RDD to sketch
   * @param sampleSizePerPartition max sample size per partition
   * @return (total number of items, an array of (partitionId, number of items, sample))
   */
  def sketch[K : ClassTag](
      rdd: RDD[K],
      sampleSizePerPartition: Int): (Long, Array[(Int, Long, Array[K])]) = {
    val shift = rdd.id
    // val classTagK = classTag[K] // to avoid serializing the entire partitioner object
    val sketched = rdd.mapPartitionsWithIndex { (idx, iter) =>
      val seed = byteswap32(idx ^ (shift << 16))
      val (sample, n) = SamplingUtils.reservoirSampleAndCount(
        iter, sampleSizePerPartition, seed)
      Iterator((idx, n, sample))
    }.collect()
    val numItems = sketched.map(_._2).sum
    (numItems, sketched)
  }

在這裡調用了RDD的collect action算子出發了作業的運作，其實此處的collection是對key進行采樣已确認key的分布情況，總之是在為做全局排序做準備。

想知道詳細的請檢視spark源碼執行吧

spark core sortBy和sortByKey探索

繼續閱讀

【51CTO學院三周年】自學路上的伴侶

線上教育巨頭多鄰國Duolingo入華一周年，中國市場馬力全開

【分類算法】什麼是分類算法定義分類與聚類分類過程方法

申請評分模型拒絕推斷（RI）方法申請評分模型拒絕推斷（RI）方法

Sql優化一：sql語句優化

Nacos 2.0 更新前後性能對比壓測

尚矽谷—韓順平—圖解 Java設計模式（結構型）（55～）

Storm編譯打包過程中遇到的一些問題及解決方法

MapReduce的幾個企業級經典面試案例MapReduce的幾個企業級經典面試案例

9.spark Core 進階2--Cashe

大資料排錯SparkSpark叢集啟動時候，JAVA_HOME is not sethadoop叢集，某台伺服器jps無任何輸出IDEAkafkahadoopspark sqlfile permissionsIDEA本地測試 - OutOfMemoryError: GC overhead limit exceededhdfs負載均衡

淺談企業活動中進行資料分析的重要性

Ambari介紹和架構原理

spark/scala關于【資源檔案】加載方法概述外部檔案加載方案測試資源檔案打包入jar包中小結

NOSQL安全攻擊

win10本地scala和spark安裝安裝scala安裝spark