[SPARK-34954][SQL] Use zstd codec name in ORC file names

### What changes were proposed in this pull request? This PR aims to add `zstd` codec names in the Spark generated ORC file names for consistency. ### Why are the changes needed? Like the other ORC supported codecs, we had better have `zstd` in the Spark generated ORC file names. Please note that there is no problem at reading/writing ORC zstd files currently. This PR only aims to revise the file name format for consistency. **SNAPPY** ``` scala> spark.range(10).repartition(1).write.option("compression", "snappy").orc("/tmp/snappy") $ ls -al /tmp/snappy total 24 drwxr-xr-x 6 dongjoon wheel 192 Apr 4 12:17 . drwxrwxrwt 14 root wheel 448 Apr 4 12:17 .. -rw-r--r-- 1 dongjoon wheel 8 Apr 4 12:17 ._SUCCESS.crc -rw-r--r-- 1 dongjoon wheel 12 Apr 4 12:17 .part-00000-833bb7ad-d1e1-48cc-9719-07b2d594aa4c-c000.snappy.orc.crc -rw-r--r-- 1 dongjoon wheel 0 Apr 4 12:17 _SUCCESS -rw-r--r-- 1 dongjoon wheel 231 Apr 4 12:17 part-00000-833bb7ad-d1e1-48cc-9719-07b2d594aa4c-c000.snappy.orc ``` **ZSTD (AS-IS)** ``` scala> spark.range(10).repartition(1).write.option("compression", "zstd").orc("/tmp/zstd") $ ls -al /tmp/zstd total 24 drwxr-xr-x 6 dongjoon wheel 192 Apr 4 12:17 . drwxrwxrwt 14 root wheel 448 Apr 4 12:17 .. -rw-r--r-- 1 dongjoon wheel 8 Apr 4 12:17 ._SUCCESS.crc -rw-r--r-- 1 dongjoon wheel 12 Apr 4 12:17 .part-00000-2f403ce9-7314-4db5-bca3-b1c1dd83335f-c000.orc.crc -rw-r--r-- 1 dongjoon wheel 0 Apr 4 12:17 _SUCCESS -rw-r--r-- 1 dongjoon wheel 231 Apr 4 12:17 part-00000-2f403ce9-7314-4db5-bca3-b1c1dd83335f-c000.orc ``` **ZSTD (After this PR)** ``` scala> spark.range(10).repartition(1).write.option("compression", "zstd").orc("/tmp/zstd_new") $ ls -al /tmp/zstd_new total 24 drwxr-xr-x 6 dongjoon wheel 192 Apr 4 12:28 . drwxrwxrwt 15 root wheel 480 Apr 4 12:28 .. -rw-r--r-- 1 dongjoon wheel 8 Apr 4 12:28 ._SUCCESS.crc -rw-r--r-- 1 dongjoon wheel 12 Apr 4 12:28 .part-00000-49d57329-7196-4caf-839c-4251c876e26b-c000.zstd.orc.crc -rw-r--r-- 1 dongjoon wheel 0 Apr 4 12:28 _SUCCESS -rw-r--r-- 1 dongjoon wheel 231 Apr 4 12:28 part-00000-49d57329-7196-4caf-839c-4251c876e26b-c000.zstd.orc ``` ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Pass the CIs with the updated UT. Closes #32051 from dongjoon-hyun/SPARK-34954. Authored-by: Dongjoon Hyun <dhyun@apple.com> Signed-off-by: Dongjoon Hyun <dhyun@apple.com>
2021-04-04 17:11:56 -07:00 · 2021-04-04 17:11:56 -07:00 · 748f05fca9
parent ba32b200e4
commit 748f05fca9
2 changed files with 3 additions and 0 deletions
--- a/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/orc/OrcUtils.scala
+++ b/sql/core/src/main/scala/org/apache/spark/sql/execution/datasources/orc/OrcUtils.scala
@ -45,6 +45,7 @@ object OrcUtils extends Logging {
    "NONE" -> "",
    "SNAPPY" -> ".snappy",
    "ZLIB" -> ".zlib",
+    "ZSTD" -> ".zstd",
    "LZO" -> ".lzo")

  def listOrcFiles(pathStr: String, conf: Configuration): Seq[Path] = {
--- a/sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/orc/OrcSourceSuite.scala
+++ b/sql/core/src/test/scala/org/apache/spark/sql/execution/datasources/orc/OrcSourceSuite.scala
@ -602,6 +602,8 @@ class OrcSourceSuite extends OrcSuite with SharedSparkSession {
      val path = dir.getAbsolutePath
      spark.range(3).write.option("compression", "zstd").orc(path)
      checkAnswer(spark.read.orc(path), Seq(Row(0), Row(1), Row(2)))
+      val files = OrcUtils.listOrcFiles(path, spark.sessionState.newHadoopConf())
+      assert(files.nonEmpty && files.forall(_.getName.contains("zstd")))
    }
  }