// GCP-Certified Professional Data Engineer
Hi, I'm Ranjit Ranjan.
I build the data pipelines that quietly keep analytics honest — batch and near-real-time systems on the Google Cloud Platform.
I am a GCP-Certified Professional Data Engineer who spends my days wrestling with massive datasets and making sure they play nicely together. With over 2 years of experience in the trenches of large-scale retail systems, I design, build, and manage both batch and near-real-time data pipelines on the Google Cloud Platform.
If data were a physical object, I’d be the plumber, the architect, and the traffic controller all rolled into one.
This blog serves as my “field notes”—a public space where I document the challenges I face and the solutions I build while navigating the modern data engineering ecosystem.
Current Focus
My Toolkit
My day-to-day stack revolves around solving complex data integration, transformation, and orchestration problems at retail-scale:
- Processing Engines: PySpark (My absolute favorite tool for heavy lifting)
- Data Warehousing: Google BigQuery
- Orchestration: Apache Airflow
- Message Brokers: Apache Kafka
- Cloud Infrastructure: Google Cloud Platform (GCP Professional Data Engineer)
The Philosophy
I don’t just write code that “works.” I focus on building pipelines that are:
- Idempotent: Reproducible without side effects.
- Scalable: Designed to handle data volume growth from Day 1.
- Observable: Properly monitored, logged, and easy to debug.
“A data engineer’s job isn’t to move data from A to B. It’s to ensure data arrives at B with unquestionable reliability and crystal-clear business value.”
Let’s Connect!
I am always open to discussing data architectures, big data challenges, or just geeking out over new open-source tools.
Feel free to connect with me professionally on LinkedIn.