Process massive datasets efficiently using distributed batch processing frameworks like Apache Spark.
Write a PySpark script to clean, aggregate, and transform a large log dataset, saving the output as partitioned Parquet files.