Data Engineering Professional TrainingProfessional data program

From pipeline developer to governed data product engineer

AI-Ready Data Engineering Professional Program

Build reliable batch, streaming and unstructured-data products with contracts, lineage, quality, access controls and observability for analytics and AI workloads.

View curriculum
Recommended background

Who should take this course

  • Working SQL and basic Python
  • Experience with databases, analytics, software or reporting
  • Git and command-line fundamentals
  • No production or personal data may be used in course labs
Capability outcomes

What you should be able to do

  • Translate business and AI use cases into owned data products
  • Engineer repeatable batch and streaming ingestion
  • Model analytical, event and unstructured data
  • Implement contracts, tests, lineage and incident response
  • Design identity, classification and policy controls
  • Measure freshness, reliability, performance and cost
Detailed curriculum

Six guided modules from fundamentals to applied work.

Every module combines instructor explanation, guided implementation and a practical milestone. Tools may be adapted to the sponsoring organization’s approved stack.

6Learning modules
3Guided projects
P1

Data products and platform architecture

Start with consumers, decisions and service expectations before selecting tools.

Topics covered

  • Analytics and AI workload discovery
  • Domain ownership and data-product boundaries
  • Lakehouse, warehouse and event architecture
  • SLIs, SLOs, security and cost constraints

Practical milestone

Create an architecture and service contract for an AI-ready data product.

Tools and platforms

Architecture canvas · SQL · Cloud data platform concepts

P2

Reliable ingestion and change

Handle source evolution and replay without silently corrupting downstream decisions.

Topics covered

  • Batch, CDC and event ingestion
  • Schema evolution and compatibility
  • Idempotency, replay and late data
  • Validation, quarantine and recovery

Practical milestone

Build a replayable ingestion path with schema-change tests.

Tools and platforms

Python · Kafka concepts · Object storage · Data contracts

P3

Modelling for analytics and AI

Create stable semantic structures alongside discoverable unstructured and vector-ready assets.

Topics covered

  • Dimensional and wide-table patterns
  • Open table formats and incremental models
  • Document, metadata and embedding pipelines
  • Semantic metrics and feature considerations

Practical milestone

Publish tested analytical models plus a governed document collection.

Tools and platforms

SQL · dbt concepts · Open table format · Vector store

P4

Quality, lineage and observability

Detect data incidents before consumers discover them through incorrect reports or AI responses.

Topics covered

  • Contract and transformation tests
  • Freshness, volume and distribution checks
  • Column-level lineage and impact analysis
  • Alerting, ownership and incident runbooks

Practical milestone

Break the pipeline deliberately, detect the failure and execute the recovery runbook.

Tools and platforms

Data tests · Lineage catalog · Orchestrator · Observability tooling

P5

Governance, privacy and access

Operationalize governance through metadata, policy and evidence rather than static documentation.

Topics covered

  • Classification and retention
  • Row, column and attribute-based access
  • Tokenization and privacy controls
  • Policy-as-code and audit evidence

Practical milestone

Implement a classified dataset with role-based access and an auditable policy decision.

Tools and platforms

Catalog · IAM · Data masking · Policy engine concepts

P6

Performance, FinOps and product review

Prove that the platform meets reliability and value requirements at an acceptable unit cost.

Topics covered

  • Workload isolation and query tuning
  • Storage and compute optimization
  • Cost allocation and unit economics
  • Consumer feedback and product roadmap

Practical milestone

Defend a production-readiness report with reliability, performance and cost evidence.

Tools and platforms

Warehouse or lakehouse · FinOps dashboard · Service review

Applied learning

Turn each module into practical project work.

PROJECT 01

Contract-driven batch and CDC data product

PROJECT 02

Governed analytical and unstructured knowledge platform

PROJECT 03

Observable AI-ready data platform with access and cost controls

Course completion

Projects, feedback and final review

Architecture and contract 15% · Pipelines and models 30% · Governance controls 20% · Reliability exercise 15% · Platform defense 20%

Completion depends on working project evidence and a clear explanation of the learner’s decisions. Detailed rubrics, lab access and vendor prerequisites are confirmed before enrolment.

Corporate or individual training

Confirm fit, prerequisites and the next cohort.