Trade Platform Operations Lead
cxm direct
📍 Remote🌐 Remote🕐 25d ago🔗 himalayas
Job Description
### Trade Platform Operations Lead
**Location:** Fully Remote
**Reports to:** CEO initially, transitioning to CTO
**Team:** Hands-on team lead with hiring authority
### About the Role
We are looking for a **Trade Platform Operations Lead** to take end-to-end ownership of our **MT4/MT5 platform operations**.
The current environment is outsourced and experiencing client-facing incidents with direct financial impact. Your first priority will be to establish reliable uptime and incident baselines, then build the monitoring, recovery, and operational capabilities needed to bring MT4/MT5 fully in-house.
This is a **hands-on turnaround and build role** with direct CEO visibility and significant autonomy.
### Key Responsibilities
* Take ownership of **MT4/MT5 infrastructure, stability, uptime, and performance**.
* Audit the existing setup and establish reliable uptime, incident, and MTTR baselines.
* Build monitoring and observability using **Grafana, Prometheus, Datadog, or equivalent**.
* Develop incident-response, recovery, and preventive-maintenance playbooks.
* Lead critical production incidents and root-cause analysis.
* Reduce client-facing incidents and improve platform reliability.
* Build and manage a small **24/7 platform operations team**.
* Hire, mentor, and establish timezone coverage and escalation procedures.
* Manage relationships with **MetaQuotes, Microsoft, and infrastructure vendors**.
* Support network-access and infrastructure initiatives required for independent platform operations.
* Report reliability metrics, risks, and progress directly to senior leadership.
* Transition MT4/MT5 operations from outsourced to fully internal ownership.
### Requirements
### Requirements
* Hands-on experience operating **MT4/MT5 infrastructure** for a broker, prop firm, or similar trading environment.
* Strong production troubleshooting and incident-management experience.
* Experience with monitoring/observability tools such as **Grafana, Prometheus, Datadog, Zabbix, or ELK**.
* Experience creating runbooks, recovery procedures, and preventive-maintenance processes.
* Experience managing or leading a technical operations team, ideally with **24/7 coverage**.
* Strong infrastructure, networking, and server knowledge.
* Experience working with external infrastructure/platform vendors.
* Ability to work independently and make decisions during high-severity incidents.
* Strong communication skills and comfort reporting directly to executives.
**Nice to have:** Forex, CFD, crypto, equities, FIX, liquidity-provider, trading gateway, or broker infrastructure experience.
### Success in the First Year
* Establish the first reliable MT4/MT5 uptime and incident baseline.
* Significantly reduce client-facing incidents and MTTR.
* Implement comprehensive monitoring and recovery playbooks.
* Build a reliable 24/7 operations team.
* Reduce dependence on outsourced platform operations.
* Establish sustainable vendor and infrastructure management.
### Benefits
### Why Join
* Direct CEO access and real operational authority.
* Build the function from the ground up.
* Fully remote with no timezone restrictions.
* Performance bonus linked to measurable uptime and incident improvements.
* Opportunity to become the long-term head of platform operations.
**Interested? Tell us about your current MT4/MT5 setup and the biggest reliability problem you've had to solve.**
Originally posted on [Himalayas](https://himalayas.app)