Employment type
Contract
Industry
Banking
Area
Incident Management, Application Monitoring
Location
Luxembourg
Remote from abroad?
No
Home office?
40%
Contract duration
until 31/07/2027 with option for extension
01
Tasks and responsibilities
- End-to-end ownership of incident lifecycle: identification, logging, categorization, prioritisation, escalation, resolution, closure, and post-incident review.
- Design and enforcement of consistent SLAs/OLAs across technology teams and third parties
- Implementation of effective major incident management protocols, including war-room activation, bridging, and crisis communication.
- Integration of incident data with problem management and change control to reduce recurrence.
- Development of clear, timely, and non-technical incident updates for business stakeholders and senior management.
- Establishment of standardised executive incident reports containing impact assessments, timeline breakdowns, root causes, and action plans.
- Creation of monthly or quarterly service health dashboards shared with business unit heads to improve transparency and trust.
- Use of communication channels (e.g., MS Teams alerts, email bulletins, intranet portals) tailored to different audiences.
- Introduction of structured follow-ups with impacted departments to rebuild confidence after major disruptions.
- Redesigned incident reporting structure at a European wealth manager, introducing weekly service disruption summaries distributed to C-suite and board members.
- Automated incident briefing packs using Power BI and ServiceNow integrations, cutting report preparation time by 70%.
- Introduced a central Major Incident Register with trend analytics, enabling proactive mitigation of recurring failure patterns.
- Created trend analysis and dashboard and committees to challenge L2 and L3 teams to continuously improve.
- Led training and simulation exercises (tabletop drills) to prepare tech and business leads for high-severity scenarios.
- End-to-end observability across: Frontend applications (web, desktop, mobile interfaces)
- Application servers, microservices, etc.
- Middleware components
- Scheduled batch jobs (EOD, month-end, etc)
- File-based integrations
- API-to-API interactions
- Downstream outputs (reports, payment instructions, regulatory filings, client statements)
- Detection of missing or delayed file arrivals
- Validation of file content integrity: size, record count, schema conformance, absence of error flags.
- Tracking of batch job executions, including start/end times, return codes, exception logs, and dependencies.
- Observation of integration touchpoints, ensuring successful message transmission and acknowledgment
- Automated reconciliation checks between source and destination systems (e.g., number of transactions sent vs. processed).
- Smart alerting on logical anomalies, such as zero-balance positions, duplicate payments, or mismatched currency conversions.
- Deployment and configuration of enterprise-grade APM tools (Dynatrace, Datadog, AppDynamics, Splunk) to correlate events across systems.
- Instrumentation of scripts and backend processes to emit custom metrics and traces.
- Use of synthetic monitoring and heartbeat probes to simulate critical workflows.
- Suppression of noise through intelligent alert grouping, deduplication, and dynamic thresholds.
- Seamless integration with incident management platforms (e.g., ServiceNow) for auto-ticket creation
- At a private bank, designed a monitoring layer for transmissions and triggers automated retries or notifications — reducing undetected payment errors.
- Implemented integrity checks using monitoring, flagging malformed position files before downstream reporting jobs began.
- Created an integration health dashboard showing real-time status APIs and file feeds connecting portfolio management, custody, and billing systems used operationally by L1/L2 support.
- Automated reconciliation between outbound transaction batches and upstream confirmation receipts, identifying mismatches before settlement cut-off.
02
Must-have criteria
- Deep operational knowledge of end-to-end incident management processes
- ITIL v3/v4
- Strong technical acumen in building holistic monitoring ecosystems
- Experience with APM tools (Dynatrace, Datadog, AppDynamics, Splunk)
- Good experience in working and navigation in a multi-national, global organisation
- Pro-active thinking and knowledge sharing
- Team player spirit
- Good communication skills
- Ability to work under pressure
- Maintains strong attention to detail in high-pressure situations
03
Nice-to-have criteria
- German or French language skills are an advantage.
04
Contract duration
Start date: asap
End date: 31/07/2027 with option for extension
05
Language requirements
Excellent English language skills
06
Application form