• in
  • git
  • smtp
  • tel

Streaming-first data platforms

Where raw events become real-time intelligence.

I design streaming-first platforms that ingest events, transform them in flight, and land clean, query-ready data across Snowflake, AWS, and Databricks — turning fragile ETL into pipelines teams can trust.

Years building pipelines
0+

Years building pipelines

Cloud certifications
0

Cloud certifications

Data & cloud platforms
0

Data & cloud platforms

Stream Monitor
prod
SRC
STREAM
XFORM
SINK
Throughputevents / sec
--:--:--ingest1,240 events → kafka
--:--:--streamwindow agg · 5s tumbling
--:--:--xformdedup · enrich · validate

12,400

events/sec

42ms

p95 latency

99.98%

uptime

01 — Capabilities

A full data platform, end to end.

The stack I use to move data from source to insight — languages, processing, orchestration, cloud, and warehouses.

L0
Languages
PythonPySparkSQLShell
L1
Processing
Apache SparkKafka
L2
Orchestrate & Transform
Ab InitioAirflowdbtControl Center
L3
Cloud
AWS
L4
Warehouse & Lakehouse
SnowflakeDatabricks
L5
Containers & CI/CD
DockerKubernetesGitHub Actions

02 — Certifications

AWS Certified Data Engineer badge

AWS Certified Data Engineer

Verified · Amazon Web Services

03 — Projects

Systems I've shipped.

End-to-end data platforms — built, deployed, and documented, not just diagrammed.

P01Live

FinShield

Real-time fraud operations center

FinShield fraud operations center overview with KPIs and live alert

Architecture

Kafka
PySpark
Snowflake
dbt
Streamlit

A real-time financial fraud detection platform. Transactions stream through Kafka and PySpark Structured Streaming into Snowflake, where risk is scored in-database with Snowpark and surfaced in a SOC-style dashboard that auto-refreshes every 10 seconds.

10s

Live refresh

8

Console pages

7

Snowflake features

1/sec

Txn stream

Apache KafkaPySparkSnowflakeSnowparkdbtStreamlitPlotlyPython
  • Snowpipe ingestion + Streams & Tasks for change-data-capture and automation
  • Dynamic Tables for self-maintaining, incrementally-refreshed aggregates
  • Snowpark Python UDF for in-database risk scoring, plus z-score anomaly detection
  • dbt staging → mart models transforming raw events into query-ready tables
  • SOC-style 8-page console: live alerts, KPIs, severity, city & merchant analytics
Showcases

04 — Experience

Delivered in production at Cognizant.

Production ETL I built and shipped end to end in banking & payments — real scale, real stakeholders, zero failures since deployment.

E01Production

Cross-Platform Extract & Enrichment Pipeline

BigQuery → SQL Server extract with a reconciliation-key enrichment loop

Data Engineer · Cognizant

Architecture

BigQuery (~25M)Stage in SQL ServerMint event IDRe-extract → CSVEnrichment teamParse + cleanseLoad enriched table

A downstream team needed ~13 months of transaction history enriched and loaded into a target table — but every record first had to carry a unique, stable identifier so the enriched data could be reconciled back to its source. I designed and built an Ab Initio pipeline spanning BigQuery and SQL Server to handle the full extract–enrich–load cycle.

~25M

Records

13 mo

History window

2-way

Enrich handshake

0

Prod failures

Ab InitioBigQuerySQL ServerShellCSVError/Reject handling
  • Extracted ~13 months of history (~25M records) from BigQuery into SQL Server, deliberately staging through the database so it could generate a unique, stable event ID per record — a key a direct cloud export wouldn't attach.
  • Re-extracted the key-stamped data to CSV with reference-code lookups resolved on unload, delivered to the enrichment team using the event ID as the reconciliation key.
  • Built the return-load flow — delimited-file parsing, cleansing and standardization, and load into the target enriched table — tying each enriched record back to its original by event ID.
  • Dedicated error, reject, and logging streams around every load isolate bad records into separate files instead of failing the job — production-grade for a large backfill.

Signal · Cloud-to-warehouse movement, a deliberate reconciliation-key design, a two-way enrichment handshake, and robust error/reject/logging.

Client work — generalized, no proprietary detail

$Sandbox — query the portfolio

Don't take my word for it. Query it.

This whole page is just data — so run SQL against it. Real tables, real rows, evaluated live in your browser.

portfolio.db — projects · skills · certs · stats
SELECT name, domain, status FROM projects
namedomainstatus
FinShieldfraudlive
OrderLakeecommercedeployed
LedgerFlowbankingdeployed
OmniLakelakehouseflagship

4 rows

try:

Live — GitHub

Still shipping.

Pulled live from GitHub the moment you loaded this page — not a screenshot, the real feed.

05 — About

From moving data to designing the systems that move it.

I'm a data engineer with four years at Cognizant, where I started in ETL — building and maintaining the pipelines that keep enterprise data flowing. Along the way I got obsessed with the parts most pipelines get wrong: duplicates, late data, schema drift, silent data loss, and numbers that don't reconcile across systems.

That pushed me from moving data to designing the platforms that move it — streaming-first ingestion, open lakehouse table formats, multi-cloud query engines, and the governance and cost controls that make data trustworthy. The projects above are where I've proven those patterns end to end, on real cloud, from scratch.

Streaming & CDCOpen lakehouse (Iceberg · Delta)Multi-cloud (Snowflake · AWS · Databricks)Data quality & governanceOrchestration (Airflow · dbt)FinOps & cost control

At a glance

Role
ETL Developer → Data Engineer
Experience
4 years · Cognizant
Education
B.Tech CSE · KL University (2022)
Based in
India
Open to
Senior Data Engineer roles
  1. 2022 — Present

    ETL Developer

    Cognizant Technology Solutions

  2. 2022

    B.Tech, Computer Science

    KL University · CGPA 8.25

Résumé

Everything here, on one page.

The full story distilled into a clean, ATS-ready résumé — grab it in whichever format you need.

One page · ATS-ready · updated Jul 2026

06 — Contact

Let's build the platform behind the numbers.

I'm open to Senior Data Engineer roles and always up for a conversation about streaming, lakehouse, and multi-cloud data platforms.

Available for Senior DE roles