Results-driven Data & Automation Engineer (Programmer Analyst at Cognizant) and technical product owner specializing in high-throughput data validation engines, autonomous CI/CD pipelines, and enterprise automation utilities. Technical Product Owner and Lead Developer of iMigrator scaled across enterprise migration programsβre-engineering core reconciliation pipelines with Polars LazyFrames and DuckDB for up to 3x throughput gains and zero-OOM execution on 50M+ row tables in benchmark runs. Proven leadership integrating GenAI (AWS Bedrock / Anthropic Claude) for automated failure diagnostics, PyInstaller multi-tier self-updaters, and collision-safe Jenkins automation.
π Kolkata, India
Engineered for high-throughput distributed computation, heterogeneous cloud databases, GenAI pipelines, and autonomous test harnesses.
Production-grade systems programming, AST parsing, high-concurrency multiprocessing, and Windows OS API automation.
Memory-optimized columnar computation engines replacing traditional memory-bound iterations with zero-OOM processing for 50M+ rows.
Deep daily integration of frontier AI coding tools, autonomous agents, and conversational intelligence across VS Code and terminal workflows.
Multi-platform integration adapters developed for iMigrator with native cross-database staging and temporary table joins.
High-throughput extraction adapters handling big data serializations, enterprise nested structures, and legacy enterprise formats.
Autonomous failure clustering and root-cause diagnostics integrating cloud foundational models for automated reconciliation.
Collision-safe distributed test runners, atomic self-updating desktop distribution, and robust execution telemetry.
Enterprise QA orchestration, automated defect lifecycle management, DDL schema drift tracking, and Informatica pipelines.
Proven product ownership, architectural migrations, and production-grade engineering at Cognizant.
psutil-based dynamic memory allocation (sizing to 70β85% of available RAM) and disk-backed chunk streaming (fetch_to_disk), enabling low-memory systems (8GB/16GB RAM) to process 50M+ records across 50 columns in 10β15 minutes in high-volume benchmark runs with zero out-of-memory errors.Launcher → Updater → Production Runtime) using PyInstaller, Windows API process handling, and enterprise network distribution with rollback protection and zero console flashing.Core platforms, autonomous assistants, and enterprise distribution systems engineered for high scale.
High-throughput enterprise data reconciliation platform adopted across large-scale migration programs. Re-engineered core pipeline with Polars LazyFrames and DuckDB, delivering up to 3x processing speed and zero-OOM execution on 50M+ row tables in benchmark runs.
β’ Enterprise Adoption & Extensibility: Scaled platform across large-scale enterprise client engagements, expanding validation coverage to 35+ automated verification rules across 10+ heterogeneous connectors.
β’ In-Warehouse Pushdown Reconciliation: Staged heterogeneous source data into temporary Snowflake tables, executing distributed joins directly in-warehouse to eliminate multi-terabyte network data transfers.
β’ Adaptive Memory Allocation: Utilized psutil to evaluate free host RAM at runtime, scaling batch chunks between 70β85% memory capacity with disk-backed chunk streaming (fetch_to_disk).
β’ Keyless Fallback & Safety: Built automated MD5 hash keyless fallback reconciliation, extraction adapters for MongoDB/DynamoDB document schemas, and automated DML execution blockers to prevent accidental destructive SQL operations.
β’ GenAI Integration: Decoupled AWS Bedrock / Claude Sonnet diagnostics with pattern clustering, sampling 3 mismatch records per pattern and returning structured JSON root-cause classifications.
Autonomous multi-modal data engineering agent built with the Google ADK during Google's 4-Hour "Build with Gemini" hackathon, earning the official Google Developer Badge & Credly Certification.
β’ Dynamic Schema Catalog: Automated Firestore schema registration cataloging database schemas, primary keys, and data types across heterogeneous sources (list_tables, get_table_details, add_table).
β’ Anti-Pattern Detection: AST parsing with sqlparse to detect full table scans, missing filters, and uncapped sorting with Snowflake, BigQuery, and PostgreSQL optimizations.
β’ Multi-Modal Architecture Generation: Produced visual ER diagrams and animated Kafka event-streaming architecture diagrams leveraging Gemini multimodal models on Vertex AI with GCS storage.
β’ Vertex AI Memory Bank: Integrated PreloadMemoryTool for session-level dialect persistence, paired with AgentEngineSandboxCodeExecutor and a responsive FastAPI/A2UI card interface.
Architected a zero-downtime, atomic client updater (Launcher → Updater → Production Runtime) distributing updates across enterprise client installations with delta synchronization.
β’ Atomic Hot-Swap: Decoupled process execution so the updater replaces running binaries without file-lock collisions, backed by rollback protection.
β’ Zero Console Flashing: Leveraged Windows API process handling to provide silent execution with clean SQLite telemetry logging and network distribution.
β’ CLI Maintenance Utility: Authored automated CLI diagnostics and repair utilities for rapid environment health checks and client distribution audits.
Production ETL integration pipeline processing fixed-width and comma-delimited healthcare feeds into dimensional Oracle tables with automated XML welcome letter distribution.
β’ Multi-Feed Processing: Cleaned, validated, and normalized multi-tier patient datasets across complex business transformations.
β’ Schema Compliance: Validated outgoing XML welcome records against strict enterprise XSD schema definitions.
Verified industry credentials, certifications, and academic foundations.
Specializing in Data Engineering, High-Throughput Reconciliation, ETL Automation, and AI-Driven Data Systems. Let's discuss modern data pipelines, Polars, DuckDB, or generative AI architecture.