Warm tip: This article is reproduced from serverfault.com, please click

scala

scala project export files

发布于 2020-11-28 23:50:06

I have the below query

val CubeData = spark.sql (""" SELECT gender, department, count(bibno) AS count FROM borrowersTable, loansTable  WHERE borrowersTable.bid = loansTable.bid GROUP BY gender,department WITH CUBE ORDER BY gender,department """)

And i want to export 4 files with specific data and names.

File1 consist of gender and departments and the name of file is geneder_departments File2 gender,null name of file is gender_null File3 departments,null name of file is departments_null File4 null,null name of file is null_null theses files are results from sql query (with cube)

i try the below

val df1 = CubeData.withColumn("combination",concat(col("gender") ,lit(","), col("department")))
df1.coalesce(1).write.partitionBy("combination").format("csv").option("header", "true").mode("overwrite").save("final")

but i took more than 4 files - combination of gender - departments. Also names of those files are random. Is it possible to choose the name of those files?

Questioner

nubie

Viewed

0

mck 2020-11-29 15:46:07

Perhaps it's a bug in Spark, I don't see any problem in your query, but the query below seems to work. You don't need to specify table names if they are unique columns.

val CubeData = spark.sql ("""
SELECT gender, department, count(bibno) AS count
FROM borrowersTable
JOIN loansTable USING(bid)
GROUP BY gender, department WITH CUBE
ORDER BY gender, department
""")

But there seems to be some problems in your file parsing, try this instead:

val borrowersDF = spark.read.format("csv").option("delimiter", "|").option("header", "True").option("inferSchema", "True").load("BORROWERS.txt")
borrowersDF.createOrReplaceTempView("borrowersTable")
val loansDF = spark.read.format("csv").option("delimiter", "|").option("header", "True").option("inferSchema", "True").load("LOANS.txt")
loansDF.createOrReplaceTempView("loansTable")

val CubeData = spark.sql ("""
SELECT gender, department, count(bibno) AS count
FROM borrowersTable
JOIN loansTable USING(bid)
GROUP BY gender, department WITH CUBE
ORDER BY gender, department
""")

热门帖子

1

推荐一些好玩的/大众的手游

2

求指教后端项目迁移方案

3

迷你洗衣机是不是都是智商税？

4

求助一个排查了半年没解决的 MySQL order by 子句导致索引失效的问题， 500 多万条记录的小表要查快两分钟

5

个人开发了一款 WordPress 主题： iPao，集成了 AI 总结功能

6

偶然发现奇游加速器会在系统里植入根证书

7

国内有蒲公英替代品推荐吗？

8

语音助手这个东西真的会监听谈话并且上传，从而泄漏隐私吗？

9

出一些有意思的域名-明盘

10

jetbrains 全家桶升级 2024 后，在滚动代码时候感觉有点掉帧

热门github

1

A multi-platform library for OpenGL, OpenGL ES, Vulkan, window and input

2

Dev tool that writes scalable apps from scratch while the developer oversees the implementation

3

shadcn/ui, but for Svelte. ✨

4

The Python Risk Identification Tool for generative AI (PyRIT) is an open access automation framework to empower security professionals and machine learning engineers to proactively find risks in their generative AI systems.

5

Performance-portable, length-agnostic SIMD with runtime dispatch

6

ZK Credo

7

OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

8

Joplin - the secure note taking and to-do app with synchronisation capabilities for Windows, macOS, Linux, Android and iOS.

9

Mamba is a new state space model architecture showing promising performance on information-dense data such as language modeling, where previous subquadratic models fall short of Transformers. It is based on the line of progress on structured state space models, with an efficient hardware-aware design and implementation in the spirit of FlashAttention.

10

This repository contains System Design resources which are useful while preparing for interviews and learning Distributed Systems

11

Curso para aprender el lenguaje de programación Python desde cero y para principiantes. 75 clases, 37 horas en vídeo, código, proyectos y grupo de chat. Fundamentos, frontend, backend, testing, IA...

12

🎓 Path to a free self-taught education in Computer Science!

13

1️⃣🐝🏎️ The One Billion Row Challenge -- A fun exploration of how quickly 1B rows from a text file can be aggregated with Java

14

A collective list of free APIs

15

📚 Freely available programming books