PrimrIQ

Become a Data Engineer

Data Engineering Mentorship Program

PrimrIQ AI Services LLP runs this programme from Noida, India. Build the pipelines that power data teams — from advanced SQL and Python ETL through Snowflake, dbt, PySpark, Airflow, Kafka, and Docker. 12 modules covering the complete modern data stack.

8 months13 Projects + 2 CapstonesMentor-reviewed

Tools covered

SQLPythonSnowflakedbtPySparkAirflowKafkaDockerAWSGreat Expectations

Roles you’ll be ready for

Data EngineerAnalytics EngineerETL DeveloperCloud Data Engineer
DATA ENGINEERINGOrchestrate the nightly load.
PrimrIQLIVE55:07
airflow
Airflowupi_risk_stream
extract_s312s
load_raw41s
dbt_buildrunning
dbt_testqueued
[06:00:22] 4 of 14 OK fct_orders[06:00:24] START mart_revenue_daily▸ backfill 7 days queued
dbtJob run 4172Success
BRONZE
raw_ordersraw_stores
SILVER
stg_ordersdim_store
GOLD
fct_orders
✓unique order_id
✓relationships dim_store
○row_count within 5%
Databricksprimriq-de-2
df = spark.read.parquet( 's3://primriq/orders')agg = df.groupBy('store_id') .agg(F.sum('net_amount'))
Stage 4/6
SR-118082,40,610
SR-204331,07,880
SR-077464,18,240
Kafkaupi.transactions
THROUGHPUT10.2k/min
TOTAL LAG18,402
partoffsetlag
08,412,1902
18,409,7744
38,414,55118,402
48,408,2200
ENGINEERINGStep 4 of 6
Make the DAG idempotent, then backfill seven days and prove the counts do not double.
RDA rerun that changes the numbers is a bug, not a retry.
Submit for review →

Duration

8 months

Modules

13

Projects

13 Projects + 2 Capstones

Format

Live + labs

Level

Beginner to intermediate

The lab

This is the environment you work in.

Real consoles, a live brief, and a mentor reading what you submit. Not a video you watch.

Airflow
PrimrIQLab 08 · Orchestrate the nightly loadLIVE55:07AR
airflowdbtreset envRUNS ON
AirflowDAGs · upi_risk_streamschedule 0 6 * * *
extract_s312s
load_raw41s
dbt_buildrunning
dbt_testqueued
notifyqueued
[06:00:12] Running task dbt_build (try 1 of 3)[06:00:14] Found 14 models, 22 tests[06:00:15] 1 of 14 OK stg_orders ........ [OK 1.2s][06:00:17] 2 of 14 OK stg_stores ........ [OK 0.6s][06:00:19] 3 of 14 OK dim_store ......... [OK 0.9s][06:00:22] 4 of 14 OK fct_orders ........ [OK 2.4s][06:00:24] 5 of 14 START mart_revenue_daily▸ backfill 2026-08-16 → 2026-08-22 queued 
dbtCloud · Job run 4172Success
LINEAGE
BRONZE
raw_ordersraw_stores
SILVER
stg_ordersstg_storesdim_store
GOLD
fct_ordersmart_revenue_daily
MODELS
14
RUNTIME
2m 08s
TESTS
✓unique order_id
✓not_null store_id
✓relationships dim_store
✓accepted_values region
✓not_null net_amount
○row_count within 5%
Databricksprimriq-de-2 · 8 coresAttached
Cmd 3df = spark.read.parquet('s3://primriq/bronze/orders') .withColumn('k', F.concat('store_id', F.rand()))agg = (df.groupBy('store_id','region') .agg(F.sum('net_amount').alias('revenue')))agg.write.mode('overwrite').saveAsTable('silver.rev')
Stage 4/6
store_idregionrevenue
SR-1180North82,40,610
SR-2043East31,07,880
SR-0774West64,18,240
SR-3391South27,93,155
shuffle read 380 MB · was 4.1 GB before salting
KafkaTopics · upi.transactions6 partitions
THROUGHPUT10.2k/min
CONSUMERS3
TOTAL LAG18,402
partoffsetcommittedlag
08,412,1908,412,1882
18,409,7748,409,7704
28,411,0038,410,9967
38,414,5518,396,14918,402
48,408,2208,408,2200
58,410,6688,410,6635
{"txn":"UPI8841207","amt":1499,"vpa":"…@okhdfc"}{"txn":"UPI8841208","amt":230,"vpa":"…@ybl"}⚠ rebalance in progress · partition 3 reassigned
DATA ENGINEERINGStep 4 of 6
Make the DAG idempotent, then backfill seven days and prove the counts do not double.Submit for review
RDA rerun that changes the numbers is a bug, not a retry.
DATA ENGINEERINGOrchestrate the nightly load.
PrimrIQLIVE55:07
airflow
Airflowupi_risk_stream
extract_s312s
load_raw41s
dbt_buildrunning
dbt_testqueued
[06:00:22] 4 of 14 OK fct_orders[06:00:24] START mart_revenue_daily▸ backfill 7 days queued
dbtJob run 4172Success
BRONZE
raw_ordersraw_stores
SILVER
stg_ordersdim_store
GOLD
fct_orders
✓unique order_id
✓relationships dim_store
○row_count within 5%
Databricksprimriq-de-2
df = spark.read.parquet( 's3://primriq/orders')agg = df.groupBy('store_id') .agg(F.sum('net_amount'))
Stage 4/6
SR-118082,40,610
SR-204331,07,880
SR-077464,18,240
Kafkaupi.transactions
THROUGHPUT10.2k/min
TOTAL LAG18,402
partoffsetlag
08,412,1902
18,409,7744
38,414,55118,402
48,408,2200
ENGINEERINGStep 4 of 6
Make the DAG idempotent, then backfill seven days and prove the counts do not double.
RDA rerun that changes the numbers is a bug, not a retry.
Submit for review →
How it runs

What a week actually looks like

Practice-first is easy to claim. Here's the machinery behind it.

STEP 01

Live session

A working session with your mentor — not a recording. You ask questions as you hit them, and leave with the week's problem defined.

01

Live, not recorded

STEP 02

Lab work

You open the browser lab and work on real, messy data. No setup, no environment config. Just the problem.

02

Real, messy data

STEP 03

Submit

You push your work — notebook, query, pipeline, dashboard — for review. Every submission, every week.

03

Every week

STEP 04

Mentor review

Your mentor reads it, grades it, and tells you what a senior would have done differently. That feedback is the actual product.

04

The actual product

Curriculum

13 modules. Every one ends in a project.

Each module includes a hands-on project. The final module contains your capstone assignments.

🚀 Opening weeksOpening Weeks — SQL & Git Foundations22 topics

No prior SQL or Git assumed. These three weeks build the two non-negotiable foundations the entire track runs on — SQL from scratch and Git for team collaboration.

What a relational database is — tables, rows, columns, primary and foreign keys
Installing PostgreSQL and connecting with pgAdmin or TablePlus
SELECT, FROM, WHERE — filtering with all comparison and null operators
ORDER BY, LIMIT, DISTINCT; COUNT, SUM, AVG, MIN, MAX
GROUP BY, HAVING — and the common mistake of using WHERE when you need HAVING
Conditional aggregation — COUNT(CASE WHEN ... THEN 1 END) pattern
INNER JOIN, LEFT JOIN — understanding the ON clause, what each row represents
Multi-table joins; subqueries in WHERE; WITH clause (CTE)
CASE WHEN — conditional logic inside a query; COALESCE for null handling
Top N per group pattern, period comparison patterns
The grain concept — what does one row in this result represent?
What version control is and why data engineers cannot live without it
git init, git clone — creating and getting a repository
git add, git commit -m — staging changes and saving a snapshot
git push, git pull — sending and receiving changes from the team
.gitignore — excluding credentials, large data files, virtual environments
git checkout -b — creating and switching to a new branch
git merge — combining branches; understanding and resolving merge conflicts
Pull requests — the code review step before merging to main
The DE team workflow — no one commits directly to main
Conventional commits — feat:, fix:, chore:, docs: — writing useful commit messages
GitHub Actions first look — triggers, jobs, steps; why automated checks matter
Module 011 sample project · 18 topics

SQL Mastery

Advanced SQL beyond tutorials — window functions with ROWS/RANGE frames, recursive CTEs, query optimisation, and the dimensional modelling decisions that determine warehouse maintainability.
What you cover
All JOIN types — INNER, LEFT, RIGHT, FULL OUTER, CROSS, self join, anti-join patterns
Subqueries — scalar, correlated, derived tables, EXISTS/NOT EXISTS
UNION, UNION ALL, INTERSECT, EXCEPT; CASE WHEN — bucketing, pivoting
Conditional aggregation — SUM(CASE WHEN … THEN 1 ELSE 0 END) pattern
Advanced string, date, and null handling functions
OVER clause — PARTITION BY, ORDER BY, ROWS/RANGE BETWEEN
ROW_NUMBER(), RANK(), DENSE_RANK(), NTILE(n)
LAG() / LEAD(), FIRST_VALUE() / LAST_VALUE()
SUM() OVER — running totals; AVG() OVER — moving averages
Percentage of total; difference from previous row
Multiple CTEs — chaining transformations; recursive CTEs — date spines, hierarchies
EXPLAIN and EXPLAIN ANALYSE — reading plan nodes, actual vs estimated row counts
Indexes — B-tree, Hash, Composite; when not to index
Partitioning, materialised views, query tuning patterns
Normalisation — 1NF, 2NF, 3NF; denormalisation for analytical workloads
OLTP vs OLAP — row-oriented vs column-oriented storage
Star schema — fact table, dimension tables, surrogate keys
SCD Type 1 (overwrite), Type 2 (valid_from/valid_to, is_current), Type 3
Module 021 sample project · 16 topics

Python for Data Engineering

Python in data engineering means robustness, error handling, and testability — not interactivity.
What you cover
Virtual environments, pip, requirements.txt, pyproject.toml
File handling — CSV, JSON, Parquet (pyarrow.parquet), Avro (fastavro)
APIs in Python — requests, authentication, pagination, rate limiting, retry with backoff
Environment variables — python-dotenv, os.environ, never hardcoding credentials
Logging — logging module, levels (DEBUG/INFO/WARNING/ERROR/CRITICAL), structured JSON logging
Exception handling — try/except/finally, custom exceptions, error propagation in pipelines
Type hints — annotating function signatures, Optional, Union, Dict, List, dataclasses
Testing pipeline code — pytest, fixtures, mocking external services with unittest.mock
Large file reading — chunksize parameter, processing in batches, memory management
Data type optimisation — downcasting integers and floats, category dtype
GroupBy performance — agg() over apply(); vectorised operations
Writing to destinations — to_csv, to_parquet, to_sql with SQLAlchemy
Schema validation — validating column names, types, null rates before writing
SQLAlchemy — create_engine(), connection strings, execute(), text()
psycopg2 — cursor, fetchall, fetchone, executemany; transactions
Connection pooling — pool_size, max_overflow; bulk inserts — COPY command
Module 031 sample project · 10 topics

Data Modelling & Warehouse Design

Schema decisions made in year one persist for five years — star schema, medallion architecture, and SCD handling are covered in depth.
What you cover
What a data warehouse is; ELT vs ETL; data lake vs lakehouse
Medallion architecture — Bronze (raw), Silver (cleaned), Gold (aggregated business layer)
Column-oriented storage — why Snowflake, BigQuery, Redshift store by column
Fact table — measures, foreign keys to dimensions, grain definition
Dimension table — descriptive attributes, surrogate key, natural key
Grain — what one row in the fact table represents, the most important decision in schema design
Conformed dimensions, degenerate dimensions, junk dimensions, role-playing dimensions
SCD Type 1 (overwrite), Type 2 (valid_from/valid_to, is_current), Type 3 (limited history)
Data Vault — Hub (business keys), Link (relationships), Satellite (attributes + history)
Schema evolution — adding nullable columns, column renaming strategies
Module 041 sample project · 11 topics

Cloud Infrastructure & AWS

The AWS services that appear in every Indian data engineering job description — IAM, S3, Glue, Athena, Lambda, and event-driven pipeline triggers.
What you cover
IAM — users, groups, roles, identity vs resource-based policies, least privilege principle
S3 — buckets, objects, prefixes, storage classes (Standard, IA, Glacier), versioning, lifecycle rules
EC2 — instance types, AMIs, security groups; VPC — subnets, route tables, NAT gateway
RDS — managed databases, Multi-AZ, read replicas; AWS CLI — aws configure, s3 cp/sync/ls
AWS Glue — serverless ETL, DynamicFrames, crawlers, Data Catalog, job bookmarks
AWS Lambda — serverless functions, event triggers (S3, SQS, EventBridge)
Amazon Athena — serverless SQL on S3; partitioning and Parquet for cost reduction
Amazon Redshift — columnar warehouse, COPY command, distribution styles
AWS Step Functions — state machine orchestration
Amazon SQS — message queuing; EventBridge — event bus, scheduling
Cost management — S3 storage pricing, Athena query costs
Module 051 sample project · 15 topics

Snowflake

Snowflake has the broadest Indian enterprise adoption among cloud warehouses — separation of compute and storage is the architectural principle every DE must understand.
What you cover
Separation of storage and compute — unlimited storage, independently scalable compute
Virtual warehouses — sizes (X-Small to 6X-Large); auto-suspend and auto-resume
Query result cache — free query reuse within 24 hours
Time Travel — querying historical data, retention period configuration
Zero-copy cloning — instant copy without duplicating storage
FLATTEN, LATERAL FLATTEN — unnesting semi-structured JSON stored as VARIANT
PARSE_JSON, ARRAY_AGG, OBJECT_CONSTRUCT — working with semi-structured data
MERGE statement — upsert pattern, WHEN MATCHED, WHEN NOT MATCHED
COPY INTO — loading from S3, Azure Blob, GCS; ON_ERROR handling
Role hierarchy — ACCOUNTADMIN, SYSADMIN, SECURITYADMIN, custom roles
Privilege management — GRANT ON, REVOKE, future grants
Resource monitors — credit limits, suspension thresholds
Snowpipe — continuous micro-batch loading from S3, auto-ingest
Dynamic data masking and row-level security — role-based data visibility
Snowflake Tasks — scheduled SQL execution, tree of tasks
Module 061 sample project · 17 topics

dbt (data build tool)

dbt transformed SQL transformations from undocumented scripts into tested, documented, version-controlled code — now in every modern data stack job description.
What you cover
What dbt does — the T in ELT, runs SQL models against the warehouse
dbt project structure — dbt_project.yml, models/, tests/, macros/, seeds/, snapshots/
Materialisation types — view (default), table, incremental, ephemeral
Incremental models — is_incremental() macro, unique_key, filtering new rows only
ref() and source() functions — building the DAG automatically from model references
Seeds — CSV files loaded as tables; Snapshots — SCD Type 2 history automated
Generic tests — not_null, unique, accepted_values, relationships
Singular tests — custom .sql files returning rows that fail the test
schema.yml — model and column descriptions, test definitions
Test severity — warn vs error; dbt-expectations package — statistical tests
dbt docs generate and dbt docs serve — documentation site, lineage DAG visualisation
Macros — Jinja2 templating, reusable SQL snippets, macro arguments
Variables — dbt_project.yml variables, --vars flag override, var() function
env_var() function — secrets and config from environment
dbt Cloud — scheduled jobs, CI/CD integration, Slim CI (only changed models)
Hooks — pre-hook and post-hook SQL; Exposures — documenting downstream dependencies
Packages — dbt_utils, dbt_date — installing via packages.yml
Module 071 sample project · 18 topics

Distributed Processing with PySpark

PySpark is required for data that does not fit on one machine — Indian companies in fintech, e-commerce, and telecom process billions of rows daily.
What you cover
Why Spark — processing data that does not fit in a single machine's memory
Driver vs Executors — driver coordinates, executors process partitions
DataFrame API — structured data, schema, Catalyst optimiser, lazy evaluation
DAG execution — jobs, stages, tasks, shuffle as the most expensive operation
SparkSession — spark = SparkSession.builder.appName("name").getOrCreate()
Reading data — spark.read.csv(), spark.read.parquet(), spark.read.json(), explicit schema
Column operations — select(), withColumn(), drop(), alias(), cast()
Filtering, aggregations — groupBy().agg(), sum(), avg(), count(), countDistinct()
Window functions — Window.partitionBy().orderBy(), row_number(), rank(), lag(), lead()
Joins — join(df2, on, how), broadcast join hint; handling nulls
UDFs and Pandas UDFs (Vectorised) — @pandas_udf
Writing data — write.mode("overwrite"), write.parquet(), partitionBy()
Partitioning strategy — repartition() vs coalesce()
Data skew — salting technique to balance partitions
Caching — cache() vs persist(), storage levels
Spark UI — jobs, stages, tasks, executor metrics, identifying bottlenecks
Adaptive Query Execution (AQE) — dynamic repartitioning, skew join optimisation
Delta Lake — ACID transactions on Parquet files, time travel, schema enforcement, merge
Module 081 sample project · 15 topics

Workflow Orchestration with Airflow

Airflow is the orchestration standard in Indian data teams — a pipeline without orchestration is a script. Airflow makes it managed, monitored, and maintainable.
What you cover
Why orchestration — managing dependencies, retries, scheduling, monitoring
Airflow components — Webserver (UI), Scheduler, Executor, Metadata DB
Executor types — LocalExecutor, CeleryExecutor, KubernetesExecutor
@dag and @task decorators — dag_id, schedule, start_date, catchup, TaskFlow API
Task dependencies — task1 >> task2, [t1, t2] >> t3; dynamic task mapping — .expand()
XCom — cross-communication between tasks, xcom_push, xcom_pull
Operators — PythonOperator, BashOperator, SnowflakeOperator, S3KeySensor
SparkSubmitOperator, DbtCloudRunJobOperator, HttpOperator
Hooks and Connections — storing credentials, conn_id reference
Idempotency — re-running a DAG produces the same result, critical for backfills
Backfilling — running historical DAG runs; Catchup=True
Branching — BranchPythonOperator, conditional task execution
Retry logic — retries parameter, retry_delay, exponential backoff
SLA callbacks — triggering alerts on SLA breaches
Pools — limiting concurrency for resource-constrained operations
Module 091 sample project · 17 topics

Streaming with Kafka

Real-time data processing is the frontier of data engineering. Kafka is the dominant message bus in Indian fintech and e-commerce.
What you cover
Why streaming — real-time processing, event-driven architectures, decoupling
Event, Topic, Partition, Offset — the four core Kafka concepts
Broker — Kafka server; ZooKeeper vs KRaft; replication — leader + followers, ISR
Producer — partition assignment strategy (round-robin, key-based)
Consumer — consumer group, partition assignment, offset management
Exactly-once semantics — idempotent producers + transactional consumers
confluent-kafka — Producer(), produce(), flush(), delivery report callback
Consumer — subscribe(), poll(), commit(), close()
Serialisation — JSON, Avro with Schema Registry — centralised schema store
Error handling — KafkaError, retriable vs fatal errors, dead letter queue pattern
ksqlDB — CREATE STREAM, CREATE TABLE, CSAS, CTAS
Windowing — tumbling windows (fixed), hopping windows (overlapping), session windows
Stream-table join — enriching events with dimension data
CDC (Change Data Capture) — Debezium connector
Kafka Connect — JDBC source, S3 sink, Snowflake sink connectors
Lambda vs Kappa architecture — design tradeoffs
Watermarks — handling late-arriving events in stream processing
Module 101 sample project · 15 topics

Containerisation & Docker

Docker is how production data pipelines are packaged and deployed — eliminating environment inconsistency across local, staging, and production.
What you cover
What Docker solves — environment consistency, dependency isolation, reproducible deployments
Image vs container — image is the blueprint, container is the running instance
Dockerfile — FROM, WORKDIR, COPY, RUN, ENV, EXPOSE, ENTRYPOINT, CMD
.dockerignore — excluding venv, .git, *.pyc, data/ from build context
docker build, docker run — port mapping (-p), volume mounting (-v), env vars (-e)
docker exec, docker logs, docker ps, docker stop, docker rm
Image layers — layer caching, ordering Dockerfile instructions for build efficiency
Multi-stage builds — separate build environment from runtime image
Packaging Python ETL scripts — COPY requirements.txt first, then source
Secrets management — AWS Secrets Manager, HashiCorp Vault, never in Dockerfile
Airflow in Docker — apache/airflow official image, volumes for dags/ and plugins/
Running dbt in Docker — dbt-core image, profiles.yml as volume mount
docker-compose.yml — services, volumes, networks, depends_on, health checks
Docker Hub, AWS ECR, GitHub Container Registry
Kubernetes basics — pods, deployments, services, ConfigMaps, Secrets; Helm
Module 111 sample project · 17 topics

CI/CD & DataOps

Data pipelines without automated testing fail silently. CI/CD for data is what makes a pipeline production-grade — not just functional.
What you cover
What CI/CD means for data — automating testing, linting, and deployment of pipeline code
GitHub Actions — .github/workflows/, triggers (push, pull_request, schedule), jobs, steps
Testing in CI — pytest runs on every push, failing tests block merges
dbt CI — dbt compile on every PR, dbt test on changed models (dbt Slim CI)
Docker CI — building and pushing images on merge to main
Secrets in CI — GitHub Actions secrets, injecting as environment variables
Deployment automation — SSH to server, pull latest image, restart container
Environment promotion — dev → staging → production with approval gates
Why data quality fails silently — no compilation errors, wrong numbers look like correct numbers
Great Expectations — ExpectationSuite, Validator, DataContext, Checkpoints
Expectation types — not_be_null, values_to_be_between, row_count_to_be_between
Running Great Expectations in Airflow — GreatExpectationsOperator, failing the DAG
Soda Core — alternative to Great Expectations, YAML-based checks
dbt tests as data quality — not_null, unique, accepted_values, relationships, custom tests
Data observability — automated anomaly detection, elementary-data
Alerting — PagerDuty/Slack integration for pipeline failures and quality check violations
Data catalogues — DataHub, Apache Atlas — lineage, ownership, quality scores
Module 121 sample project · 4 topics

Capstone Projects

Two production-grade pipeline systems — fully tested, documented, and deployable.
What you cover
Capstone 1 — From Raw S3 to CFO Dashboard — A Production Medallion Pipeline for an Indian E-commerce Platform
Capstone 2 — 10,000 Transactions Per Minute — Real-Time UPI Risk and Analytics Pipeline for a Fintech Startup
Both capstones demonstrated running end-to-end
Architecture diagrams and written design decisions for both capstones
Outcomes

What you will learn

01Write production-grade SQL — window functions, recursive CTEs, optimisation for billion-row tables
02Build robust Python ETL pipelines with error handling, logging, and pytest test suites
03Design star schema dimensional models with SCD Type 2 in Snowflake
04Build dbt projects with tested, documented, version-controlled transformation models
05Process large datasets with PySpark DataFrames, Delta Lake, and Spark UI tuning
06Orchestrate pipelines with Airflow DAGs, idempotent design, and automated alerting
07Build real-time event processing pipelines with Kafka, ksqlDB, and Kafka Connect
08Package and deploy pipelines with Docker, CI/CD, and automated Great Expectations checks
Straight answers

The questions people actually ask

Do I need cloud experience before joining?

No. Module 4 covers AWS from IAM fundamentals through Glue, Athena, and S3. A free-tier AWS account is sufficient for all projects in that module.

Do I need to pay separately for any tools or cloud services?

No. All tool access and platform costs are covered as part of your course fee. You will not need to set up billing or pay for any external services — just show up and build.

Is Kafka difficult to set up?

Module 9 uses Kafka running in Docker Compose, which requires no manual installation. The focus is on writing producers, consumers, and ksqlDB queries — not cluster administration.

Will I be able to work with dbt at a company after this track?

Yes. Module 6 covers the full dbt workflow including incremental models, tests, macros, docs, and CI — the same workflow used at Indian companies running dbt in production.

Is there a refund policy?

Yes. Attend the onboarding session and apply for a full refund within 7 days of purchase if you are not satisfied. No questions asked.

Questions

Frequently asked

Anything else, write to info@primriq.com — a human replies.

Do I need to know how to code before I start?

Every track opens with four weeks of Python and SQL foundations that assume nothing. If you have written code before, those weeks go quickly; if you have not, they are the reason the rest of the programme is reachable.

How much time should I plan for each week?

Eight to ten hours: one live session, lab work on your own schedule, and a submission your mentor reviews. Labs stay open, so a heavy week at college or work does not reset your progress.

What happens if I am not satisfied?

Attend, try the labs, and if it is not what you expected, apply for a full refund within seven days of purchase. No conditions beyond that.

Is the certificate worth anything to an employer?

The certificate is verifiable and states what you completed. What carries weight in an interview is the portfolio behind it — which is why every module ends in a project you can walk a hiring manager through line by line.

Do you guarantee a job at the end?

No, and we will not pretend otherwise. You get resume work, mock interviews and career guidance, plus real projects you can talk through in detail. Interviews are still yours to win.

Course fee

₹17,999 + GST

One-time · includes all projects, live sessions and mentor reviews

Attended and not satisfied? Apply for a full refund within 7 days of purchase.

Industry range

₹5–10 LPA

Fresher, India — market benchmark, not a guarantee

Ready to start your Data Engineering journey?

13 Projects + 2 Capstones to build, a mentor on every submission, and 8 months of live sessions.

See the curriculum

7-day full refund · No setup required · Career prep included