Job Description: DataBricks - Lead Programmer Analyst

We are looking for a highly self-motivated individual with DataBricks development as a Lead Programmer Analyst:

Required Skills & Qualifications:

    • Experience should have 5 to 7 Years of Data Engineering.
    • Expert-level Apache Spark skills using PySpark (Scala a plus).
    • Strong proficiency in SQL for data transformation and performance tuning.
    • Solid experience with Delta Lake and the Lakehouse/medallion architecture.
    • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP) and its data services.
    • Experience with workflow orchestration (Databricks Workflows, Delta Live Tables, or Airflow).
    • Strong understanding of data modeling, data warehousing, and ETL/ELT design patterns.
    • Experience with Python for data engineering and automation.
    • Familiarity with CI/CD, Git, and DevOps practices for data.
    • Good understanding of data governance, security, and Unity Catalog.
    • Strong problem-solving and communication skills.
    • Bachelor's or Master's degree in Computer Science, Engineering, or a related field.

Good To Have:

    • Databricks Certified Data Engineer Associate or Professional certification.
    • Experience with streaming technologies (Structured Streaming, Kafka, Event Hubs, Pub/Sub).
    • Exposure to machine learning workflows and MLflow.
    • Experience with Infrastructure-as-Code (Terraform) and cloud cost optimization.
    • Experience working in Agile delivery environments.
    • Experience in data warehouse design and maintenance
    • Experience in agile development processes using Jira and Confluence.
    • Experience in cross-functional teams.

Key Responsibilities:

    • Design, develop, and maintain scalable ETL/ELT data pipelines on the Databricks Lakehouse Platform.
    • Build and optimize Apache Spark (PySpark/Scala) jobs for batch and streaming data processing.
    • Implement and manage Delta Lake tables, including partitioning, schema evolution, and performance tuning (Z-ordering, caching, file compaction).
    • Develop and orchestrate workflows using Databricks Workflows, Delta Live Tables (DLT), and job scheduling tools such as Airflow.
    • Implement the medallion architecture (bronze, silver, gold layers) for reliable, governed data.
    • Integrate Databricks with cloud data services (AWS, Azure, or GCP) and source systems.
    • Apply data governance and security using Unity Catalog, access controls, and lineage.
    • Optimize cluster configuration, cost, and performance across workloads.
    • Collaborate with data analysts, data scientists, and business stakeholders to deliver trusted datasets.
    • Ensure data quality, testing, monitoring, and observability across pipelines.
    • Contribute to CI/CD practices for data engineering (version control, automated deployment).