
from neuron-cli16
Expert guidance for writing high-performance Spark jobs, optimizing ETL pipelines, and tuning distributed data processing.
This skill transforms the agent into a senior Apache Spark engineer capable of architecting and optimizing large-scale distributed data processing pipelines. It provides deep expertise in PySpark and Scala, focusing on production-grade reliability and performance.
Use this skill when writing new Spark jobs, debugging performance bottlenecks (like shuffle spill or data skew), configuring cluster memory, or implementing structured streaming analytics.
collect() OOMs or inefficient UDFs.Compatible with any LLM-based agent capable of writing and executing Python/Scala code (Claude Code, Cursor, etc.).
This skill has not been reviewed by our automated audit pipeline yet.