SYSTEM STATUS: OPERATIONAL

Thejas Kumar Patel

Infrastructure & Production Support Engineer

I keep banking-scale systems observable, stable, and self-healing — turning noisy alerts into clean dashboards, and incidents into root causes fixed for good.

15+ yrs in production & infra ops
80% faster incident response via automation
99%+ SLA / KPI compliance maintained
1,500+ MQ instances managed
Monitoring & Observability
Automation & Self-Healing
Incident Response
Cloud & Middleware

01 / About

Fifteen years of keeping other people's systems up.

I'm a production support and infrastructure engineer who has spent the last decade and a half inside banking, fintech, and airline environments — the kind where downtime shows up on a balance sheet or a departures board. My work sits at the intersection of observability, automation, and incident response: building the dashboards and alerts that catch problems early, and the self-healing scripts that fix the recurring ones before anyone gets paged.

I've supported everything from WebSphere MQ messaging backbones to Azure Red Hat OpenShift clusters, and I've learned that most outages aren't mysteries — they're missing alerts, undocumented dependencies, or manual steps nobody automated yet. Closing those gaps is the job.

Consistently recognized by clients for technical ownership and being a dependable, go-to member of the team.

02 / Deployment Log

Experience

Production Support / Operations Engineer

  • Ran production support and operations for banking-scale systems, focused on stability, observability, and incident reduction.
  • Built and maintained Splunk dashboards for real-time visibility across application, infrastructure, and operational health.
  • Set up proactive alerting across Splunk, AppDynamics, Grafana, Prometheus, and Splunk On-Call.
  • Built Power Automate workflows to alert the team the moment an incident landed in our queue, cutting response time by roughly 80%.
  • Owned RCA end-to-end and automated repetitive operational tasks into self-healing, auto-triage workflows — cutting MTTR by 80% and manual triage effort by half, since triage could be decided directly from logs already surfaced rather than dug up by hand.
  • Consistently maintained >99% SLA/KPI compliance across incidents.
  • Kept Autosys-scheduled batch workloads running reliably in coordination with application teams.
SplunkAppDynamicsGrafanaPrometheusSplunk On-CallPower AutomateAutosys

Cloud / Middleware / Infrastructure Support Engineer

  • Supported BAU deployments and network connectivity across on-prem data centers, Azure Red Hat OpenShift (ARO), and Pivotal Cloud Foundry.
  • Provided on-call incident support focused on network, middleware, and cloud connectivity, coordinating resolution across teams.
  • Integrated Wiz CSPM into ARO clusters using Terraform, improving visibility into exposure and misconfigurations.
  • Provisioned Azure VMs, Cosmos DB, and ARO clusters in new landing zones.
AzureAzure Red Hat OpenShiftTerraformWiz CSPMPCFCosmos DB

Infrastructure Support Engineer

  • Delivered end-to-end environment and application support for the lower environments of a major bank.
  • Stood up monitoring, alerting, and dashboards across on-prem and cloud using Splunk, Dynatrace, and Grafana.
  • Triaged issues spanning network, MQ, HAProxy, API gateway, WAS, Tomcat, DataPower, CyberArk, ADFS, AWS, and Kubernetes.
  • Protected testing windows by ensuring environment availability after every CI/CD deployment.
  • Applied self-healing automation and coordinated response on standard and high-severity incidents.
SplunkDynatraceKibanaGrafanaWASKubernetesAWS

WebSphere MQ & WAS Administrator

  • Designed and implemented WebSphere MQ messaging solutions, including capacity tuning and security for data in transit and at rest.
  • Built high availability for MQ using OS-level and MQ clustering across Linux, Solaris, AIX, HP-UX, and OpenVMS.
  • Managed 1,500+ WebSphere MQ instances supporting cross-platform communication.
  • Handled severity incidents and RCA; ran build and deployment for Java and .NET apps on WebSphere Application Server.
WebSphere MQWASVCS

WebSphere MQ & WAS Administrator

  • Delivered 24/7 support for MQ and WAS infrastructure for the largest global airline alliance.
  • Migrated MQ 6.0 → 7.0 and WAS v6.1 → v7.0 with minimal service disruption.
  • Supported implementation planning, release execution, and post-release stabilization.
WebSphere MQWAS

03 / Stack

Tools & Technologies

Monitoring & Observability

Splunk, AppDynamics, Grafana, Prometheus, Dynatrace, Kibana, PagerDuty, Splunk On-Call

Cloud & Platforms

Azure, Azure Red Hat OpenShift, Pivotal Cloud Foundry, AWS, Kubernetes, Terraform, Wiz CSPM, CrowdStrike

Middleware & Messaging

WebSphere MQ, WebSphere Application Server, IHS, Tomcat, Apache, WMB, API Gateway, F5, HAProxy, DataPower

Automation & Ops

Autosys, self-healing & auto-triage scripting, CI/CD environment support, incident / change / problem management (JIRA, ServiceNow, HPSM)

Databases

Oracle, Microsoft SQL Server, MongoDB, DB2, Azure Cosmos DB

Languages & OS

C, C++, Python, Perl, SQL, HTML/XML · RHEL, Solaris, AIX, OpenVMS, Windows Server

04 / Certifications

Certifications

ISC2 Certified in Cybersecurity
AWS Certified Cloud Practitioner
IBM MQ V7 Administrator
IBM WAS V7 ND Administrator

05 / Education

Background

B.Tech

06 / Contact

Let's talk about keeping things running.

Open to conversations about production support, SRE, and infrastructure roles in banking, fintech, and beyond.