Results-driven Data & Automation Engineer (Programmer Analyst at Cognizant) owning and maintaining enterprise data validation platforms, CI/CD automation, and test engineering utilities. Platform Owner and Maintainer of iMigrator across multiple projectsβstabilising the inherited core engine and independently introducing DuckDB-based validation to achieve up to 3x throughput gains in benchmark runs and eliminate crashes on low-RAM systems. Co-developed GenAI failure diagnostics (AWS Bedrock / Claude) and Snowflake pushdown reconciliation, built standalone self-updaters and packaging, and improved Jenkins CI/CD automation across multiple project workflows.
π Kolkata, India
Engineered for high-throughput distributed computation, heterogeneous cloud databases, GenAI pipelines, and autonomous test harnesses.
Production-grade systems programming, AST parsing, high-concurrency multiprocessing, and Windows OS API automation.
Memory-optimized columnar computation engines replacing traditional memory-bound iterations with zero-OOM processing for 50M+ rows.
Deep daily integration of frontier AI coding tools, autonomous agents, and conversational intelligence across VS Code and terminal workflows.
Multi-platform integration adapters developed for iMigrator with native cross-database staging and temporary table joins.
High-throughput extraction adapters handling big data serializations, enterprise nested structures, and legacy enterprise formats.
Autonomous failure clustering and root-cause diagnostics integrating cloud foundational models for automated reconciliation.
Collision-safe distributed test runners, atomic self-updating desktop distribution, and robust execution telemetry.
Enterprise QA orchestration, automated defect lifecycle management, DDL schema drift tracking, and Informatica pipelines.
Proven product ownership, architectural migrations, and production-grade engineering at Cognizant.
Core platforms, autonomous assistants, and enterprise distribution systems engineered for high scale.
High-throughput enterprise data reconciliation platform adopted across multiple projects. Stabilised the inherited core pipeline and independently introduced DuckDB-based validation, delivering up to 3x processing speed in benchmark runs and eliminating crashes on low-RAM systems.
β’ Platform Ownership & Connectors: Own and maintain the platform across multiple projects; built and extended support for MongoDB, Snowflake (token-based auth), CTRL files, and DB2 while supporting users across 15+ connectors.
β’ In-Warehouse Pushdown Reconciliation: Co-developed cross-database pushdown reconciliation staging heterogeneous source data into temporary Snowflake tables for distributed in-warehouse joins.
β’ Chunked Processing & Stability: Improved and stabilised inherited chunked processing routines and dynamic memory controls, eliminating crashes on user workstations.
β’ Keyless Fallback & Safety: Owned and maintained keyless fallback reconciliation routines and automated DML execution safeguards.
β’ GenAI Failure Diagnostics: Co-developed autonomous diagnostics with AWS Bedrock / Claude Sonnet, clustering mismatch patterns and returning structured JSON root-cause classifications.
β’ Reporting & Packaging: Improved HTML validation reports with enhanced console logging and fallbacks; built standalone self-updaters and packaging via PyInstaller.
Autonomous multi-modal data engineering agent built with the Google ADK during Google's 4-Hour "Build with Gemini" hackathon, earning the official Google Developer Badge & Credly Certification.
β’ Dynamic Schema Catalog: Automated Firestore schema registration cataloging database schemas, primary keys, and data types across heterogeneous sources (list_tables, get_table_details, add_table).
β’ Anti-Pattern Detection: AST parsing with sqlparse to detect full table scans, missing filters, and uncapped sorting with Snowflake, BigQuery, and PostgreSQL optimizations.
β’ Multi-Modal Architecture Generation: Produced visual ER diagrams and animated Kafka event-streaming architecture diagrams leveraging Gemini multimodal models on Vertex AI with GCS storage.
β’ Vertex AI Memory Bank: Integrated PreloadMemoryTool for session-level dialect persistence, paired with AgentEngineSandboxCodeExecutor and a responsive FastAPI/A2UI card interface.
Architected a zero-downtime, atomic client updater (Launcher → Updater → Production Runtime) distributing updates across enterprise client installations with automated build verification and zero-downtime hot-swap.
β’ Atomic Hot-Swap: Decoupled process execution so the updater replaces running binaries without file-lock collisions, backed by rollback protection.
β’ Zero Console Flashing: Leveraged Windows API process handling to provide silent execution with clean SQLite telemetry logging and network distribution.
β’ CLI Maintenance Utility: Authored automated CLI diagnostics and repair utilities for rapid environment health checks and client distribution audits.
Production ETL integration pipeline processing fixed-width and comma-delimited healthcare feeds into dimensional Oracle tables with automated XML welcome letter distribution.
β’ Multi-Feed Processing: Cleaned, validated, and normalized multi-tier patient datasets across complex business transformations.
β’ Schema Compliance: Validated outgoing XML welcome records against strict enterprise XSD schema definitions.
Verified industry credentials, certifications, and academic foundations.
Specializing in Data Engineering, High-Throughput Reconciliation, ETL Automation, and AI-Driven Data Systems. Let's discuss modern data pipelines, Polars, DuckDB, or generative AI architecture.