Job Summary
The Sophis Platform Engineer is responsible for maintaining the stability, security, availability, and operational readiness of the Sophis infrastructure and supporting environments.
The role is primarily focused on platform lifecycle management, infrastructure maintenance, monitoring, incident resolution, patching, disaster recovery, and technical readiness. It does not require deep Sophis functional or business-product knowledge; however, a basic understanding of the Sophis application architecture and its infrastructure dependencies is expected.
Key Responsibilities
Platform Lifecycle Management Manage the end-to-end lifecycle of Sophis platform infrastructure and supporting technologies. Maintain operating systems, middleware, databases, storage, network connectivity, and other infrastructure components supporting Sophis. Track infrastructure versions, vendor support dates, end-of-life dates, and upgrade requirements. Develop and maintain lifecycle-management plans for platform components. Coordinate the replacement, upgrade, or decommissioning of obsolete and unsupported technologies. Maintain accurate platform inventories, configuration records, dependency maps, and lifecycle documentation. Planned Maintenance Plan and execute operating-system, middleware, database, security, and infrastructure patching activities. Coordinate scheduled maintenance windows with application, infrastructure, security, database, network, and business teams. Perform pre-maintenance checks, technical impact assessments, backup validation, and rollback planning. Execute infrastructure upgrades and platform refresh activities. Conduct post-maintenance validation to confirm platform availability, connectivity, capacity, and overall system health. Maintain implementation plans, validation checklists, change records, and evidence of completed activities. Ensure maintenance activities comply with organisational change-management and governance processes. Disaster Recovery and Resilience Plan, coordinate, and execute disaster recovery exercises for Sophis environments. Validate backup, restoration, replication, failover, and recovery procedures. Confirm that recovery-time and recovery-point objectives can be achieved. Identify issues discovered during disaster recovery exercises and track remediation actions through completion. Maintain disaster recovery documentation, recovery procedures, technical runbooks, and supporting evidence. Support improvements to platform resilience, redundancy, and recoverability. Roadmap and Technical Readiness Track upcoming technical milestones, infrastructure upgrades, security requirements, vendor changes, and platform dependencies. Assess the impact of technology changes on Sophis environments. Coordinate with application, database, network, storage, cloud, security, and other infrastructure teams. Ensure environments are technically ready for application releases, infrastructure changes, migrations, and business initiatives. Identify lifecycle risks, capacity constraints, unsupported components, and technical dependencies. Provide platform-readiness updates, risk assessments, and remediation recommendations. Contribute to platform roadmaps and long-term infrastructure planning. Monitoring and Incident Handling Monitor Sophis infrastructure using enterprise monitoring and alerting tools. Own and resolve infrastructure-focused incidents relating to: Memory and CPU utilisation Disk space and storage capacity Server availability and system health Network connectivity and latency Middleware and service availability Database connectivity Batch, scheduler, or service failures caused by infrastructure issues Backup, replication, and recovery failures Investigate alerts, identify root causes, and restore services within agreed service levels. Coordinate incident resolution across infrastructure, application, database, network, and vendor teams. Escalate product-specific or functional issues to the appropriate Sophis application-support team. Participate in major-incident calls and provide clear technical updates. Complete root-cause analysis and implement preventive actions for recurring platform incidents. Security and Compliance Ensure Sophis infrastructure remains secure, patched, supported, and compliant with organisational standards. Address security vulnerabilities and infrastructure-related audit findings. Support vulnerability scanning, remediation planning, evidence collection, and compliance reporting. Apply approved security hardening standards to platform components. Maintain
Skill Requirements
Required Skills and Experience
• Experience in platform engineering, infrastructure operations, production support, or systems administration.
• Strong experience with infrastructure lifecycle-management activities.
• Experience planning and executing patching, upgrades, migrations, and technology refreshes.
• Experience supporting business-critical production environments.
• Knowledge of Windows and/or Linux server administration.
• Understanding of networking, storage, databases, middleware, system monitoring, and backup technologies.
• Experience handling infrastructure incidents involving memory, CPU, storage, connectivity, availability, and system health.
• Familiarity with disaster recovery, backup validation, failover testing, and resilience exercises.
• Knowledge of ITIL processes, including incident, problem, change, configuration, and capacity management.
• Ability to produce implementation plans, rollback plans, technical runbooks, and operational documentation.
• Strong analytical, troubleshooting, coordination, and communication skills.
• Ability to work with application, infrastructure, security, database, network, cloud, and vendor teams.
• Basic knowledge of Sophis architecture, services, environments, and technical dependencies.
Desirable Skills
• Previous experience supporting Sophis or another capital-markets or trading platform.
• Experience working within banking, financial services, investment management, or capital-markets environments.
• Knowledge of Microsoft SQL Server, Oracle, or other enterprise database platforms.
• Experience with enterprise monitoring, scheduling, automation, and configuration-management tools.
• Exposure to cloud or hybrid infrastructure.
• Experience with vulnerability-management and security-compliance processes.
• Scripting or automation skills using PowerShell, Python, Bash, or similar technologies.
Other Requirements
Key Deliverables
- Stable and available Sophis environments.
- Timely completion of patching and infrastructure-maintenance activities.
- Up-to-date lifecycle and technology-obsolescence plans.
- Successful completion of disaster recovery exercises.
- Prompt resolution of infrastructure incidents and monitoring alerts.
- Accurate platform documentation, inventories, and support runbooks.
- Effective remediation of vulnerabilities and unsupported components.
- Infrastructure readiness for application releases, upgrades, and business initiatives.
- Reduced operational risk through proactive monitoring, capacity planning, and lifecycle management.
Key Stakeholders
- Sophis Application Support Team
- Infrastructure and Platform Engineering Teams
- Database Administration Team
- Network and Storage Teams
- Information Security and Cybersecurity Teams
- Service Management and Change Management
- Business Continuity and Disaster Recovery Teams
- Release and Environment Management Teams
- External technology vendors and service providers
Success Measures
- Platform availability and service stability.
- Compliance with patching and vulnerability-remediation targets.
- Reduction in infrastructure-related incidents and recurring alerts.
- Successful completion of planned maintenance and disaster recovery exercises.
- Timely remediation of lifecycle, supportability, and capacity risks.
- Adherence to incident, change, security, and operational service-level agreements.
Accuracy and completeness of platform documentation and lifecycle records.