Files
DBAdmin/inbox/Bi-Weekly Engineering & Operations Sync.md
T
Paweł Domański 3d99d13a2c Jun 2, 2026, 1:30 PM
2026-06-02 11:30:27 +00:00

16 KiB

type, title, description, tags, peopleInvolved, type_, date, dnb
type title description tags peopleInvolved type_ date dnb
Meeting Bi-Weekly Engineering & Operations Sync null
null null 2026-01-22T08:30:00.000Z -> 2026-01-22T09:30:00.000Z null

Bi-Weekly Engineering & Operations Sync

Notatka w word i muszę ją przepisać

https://mydnb-my.sharepoint.com/:fl:/g/personal/yakymyshynm_dnb_com/IQDauhw-Y3F8RYBbQFGtchWBAdRRNHsc-PiMvoiCyOBVROI?nav=cz0lMkZwZXJzb25hbCUyRnlha3lteXNoeW5tX2RuYl9jb20mZD1iIXAxVFI2QmwzdEVHLUxCOVhxWllnenJXMlZIdVEtejlNbjhZR3d0UDJCUkxyRXlJdFVpX2hSTE5aOUxTLUVsS2MmZj0wMVNETEFCTEcyWElPRDRZM1JQUkNZQVcyQUtHV1hFRk1CJmM9JTJGJmZsdWlkPTEmYT1UZWFtcyZwPSU0MGZsdWlkeCUyRmxvb3AtcGFnZS1jb250YWluZXI%3D

Meeting notes

Project updates

  • Ewa informed participants that the project team has completed their part of the Venify process, provided all necessary details, and the certificates should be replaced.

  • Ewa stated that application teams are receiving reminders to log tickets for certificate renewal, and tickets are being received as confirmation.

  • Naser confirmed that the pre-migration project is completed, with a final sync meeting scheduled with Oleksandr later in the day.

Certificate renewal process

  • The current certificate renewal process requires manual intervention because the automated flow is not yet updated, resulting in tickets being incorrectly assigned to the Windows team instead of the Venify admins.

  • Ewa explained that once the new automated flow is implemented, installation tasks will be routed to the appropriate team based on server type (Linux, Windows, or AWS).

Wildcard certificate approvals

  • Peter confirmed that wildcard certificates for internal AWS use are preapproved and do not require additional approval from Jay DePaul.

Wildcard certificate management

  • Oleksandr raised concerns that replacing wildcard certificates may lead to incidents due to lack of documentation on previous wildcard usage.

Certificate management workflows

  • Adrian stated that switching from wildcard to dedicated SSL certificates should follow a separate workflow managed by application teams, as this process may introduce risks and requires verification by the teams before replacement.

  • Jimmy explained that the policy is to install certificates on load balancers rather than servers, and that application teams should not have direct server access, aligning with the CICD module.

Automation initiatives

  • Jimmy mentioned that there is an ongoing project this year to review the possibility of integrating certificate replacement and renewal through Bamboo automation.

Load balancer migration

  • Participants discussed challenges with deploying code prepared by the AO team for load balancer migrations, highlighting concerns about responsibility and lack of application-specific knowledge.

  • Ewa explained that AO prepares code for F5 replacement because of their expertise, and the team relies on their support to facilitate the migration process.

  • Oleksandr explained that deploying code for load balancers requires deep application knowledge, so the team relies on code prepared by AO or app teams and only validates and applies it.

  • Ewa outlined three options for code deployment on Nginx: continue current process with app teams preparing code, involve the team more in projects to prepare code themselves, or allow AO team to deploy directly, noting unauthorized changes will be reported.

Oracle product review

  • Ewa informed participants that a major review of Oracle products was completed, and Dun and Brasi is not aligned with Oracle on usage.

Oracle Java compliance

  • Ewa informed participants that the legal team, together with Oracle, decided all Oracle Java installations must be uninstalled from company servers, regardless of version or previous licensing terms.

  • Ewa stated that server owners, including Leschek and possibly Adrian, have been notified to uninstall Oracle Java from their respective servers following an audit request.

  • Ewa confirmed that all application teams received instructions to uninstall Oracle Java, and any tickets requesting installation of Oracle Java should be denied.

Oracle database management

  • Ewa stated that requests for new Oracle database installations should be handled with caution, and Adrian will discuss a specific request with Beata to ensure compliance.

Server migration and support

  • Ewa explained that teams needing help with server migration or end-of-life replacements must submit requests through ServiceNow, and assistance will be coordinated by the manager.

  • Ewa explained that local market teams must log server and DNS requests in ServiceNow and drive their own migration processes, with operations support provided by the available person rather than a dedicated resource.

ServiceNow integration

  • Ewa stated that an enhancement request will be logged to obtain API access to ServiceNow, enabling better integration and automation for database and monitoring tasks.

Splunk deployment

  • Jimmy reported that the Splunk deployment project is nearing completion, with disk monitoring redesigned to use dynamic thresholds instead of macros, and the team plans to review the new approach with all operations teams.

Disk monitoring improvements

  • Jimmy demonstrated the new disk monitoring thresholds, which use dynamic levels based on drive size instead of fixed percentages.

  • Jimmy explained that the new alert system applies default thresholds to all servers, with exceptions possible, and alerts are triggered based on specific free space levels and durations.

Disk space alert routing

  • Jimmy agreed to investigate whether incidents can be configured to route to different teams based on whether the issue is with the system drive or data drive.

  • Participants discussed sending warning-level disk space alerts to server owners and critical-level alerts directly to the operations team, with ECC monitoring after hours.

Disk space alert escalation

  • Leszek raised the question of whether X Matters should be used to notify the on-call team directly for critical disk space alerts, bypassing ECC.

Disk monitoring and alerting

  • Participants discussed that disk space incidents and alerts are only generated for production systems, while non-production systems have alerts in Zabbix but do not generate incidents.

  • Jimmy explained that exceptions to the default disk monitoring rules can be managed by tagging specific disks in the file inventory, allowing for custom detectors and alert thresholds.

  • Participants agreed that incidents should be created when monitoring agents stop responding, to prevent these issues from being overlooked, and discussed whether the incident should be triggered after a certain number of hours or days.

Monitoring and server management

  • Participants agreed that incidents should be created if a monitoring agent is unresponsive for one hour on production systems.

  • Mats and Jimmy explained that detector configurations and dashboards in Splunk are managed by code and can be restored or updated automatically, with documentation available via a shared link.

Follow-up tasks

Task Assigned to Due date Bucket
Discuss certificate replacement responsibilities and global practices with Beata and Jeff (Adrian, Lesha)
Uninstall Oracle Java from all company servers as directed by the legal team (Leschek, Adrian)
Decide whether to generate tickets for warning-level disk space alerts (80%) or only for critical-level alerts (90%) (participants)
Decide whether disk space alert tickets should be sent to the server owner or directly to the operations team (participants)
Add maintenance mode step to the server decommissioning process before powering off servers (Jimmy)

Meeting notes

Project updates

  • Ewa informed participants that the project team has completed their part of the Venify process, provided all necessary details, and the certificates should be replaced.

  • Ewa stated that application teams are receiving reminders to log tickets for certificate renewal, and tickets are being received as confirmation.

  • Naser confirmed that the pre-migration project is completed, with a final sync meeting scheduled with Oleksandr later in the day.

Certificate renewal process

  • The current certificate renewal process requires manual intervention because the automated flow is not yet updated, resulting in tickets being incorrectly assigned to the Windows team instead of the Venify admins.

  • Ewa explained that once the new automated flow is implemented, installation tasks will be routed to the appropriate team based on server type (Linux, Windows, or AWS).

Wildcard certificate approvals

  • Peter confirmed that wildcard certificates for internal AWS use are preapproved and do not require additional approval from Jay DePaul.

Wildcard certificate management

  • Oleksandr raised concerns that replacing wildcard certificates may lead to incidents due to lack of documentation on previous wildcard usage.

Certificate management workflows

  • Adrian stated that switching from wildcard to dedicated SSL certificates should follow a separate workflow managed by application teams, as this process may introduce risks and requires verification by the teams before replacement.

  • Jimmy explained that the policy is to install certificates on load balancers rather than servers, and that application teams should not have direct server access, aligning with the CICD module.

Automation initiatives

  • Jimmy mentioned that there is an ongoing project this year to review the possibility of integrating certificate replacement and renewal through Bamboo automation.

Load balancer migration

  • Participants discussed challenges with deploying code prepared by the AO team for load balancer migrations, highlighting concerns about responsibility and lack of application-specific knowledge.

  • Ewa explained that AO prepares code for F5 replacement because of their expertise, and the team relies on their support to facilitate the migration process.

  • Oleksandr explained that deploying code for load balancers requires deep application knowledge, so the team relies on code prepared by AO or app teams and only validates and applies it.

  • Ewa outlined three options for code deployment on Nginx: continue current process with app teams preparing code, involve the team more in projects to prepare code themselves, or allow AO team to deploy directly, noting unauthorized changes will be reported.

Oracle product review

  • Ewa informed participants that a major review of Oracle products was completed, and Dun and Brasi is not aligned with Oracle on usage.

Oracle Java compliance

  • Ewa informed participants that the legal team, together with Oracle, decided all Oracle Java installations must be uninstalled from company servers, regardless of version or previous licensing terms.

  • Ewa stated that server owners, including Leschek and possibly Adrian, have been notified to uninstall Oracle Java from their respective servers following an audit request.

  • Ewa confirmed that all application teams received instructions to uninstall Oracle Java, and any tickets requesting installation of Oracle Java should be denied.

Oracle database management

  • Ewa stated that requests for new Oracle database installations should be handled with caution, and Adrian will discuss a specific request with Beata to ensure compliance.

Server migration and support

  • Ewa explained that teams needing help with server migration or end-of-life replacements must submit requests through ServiceNow, and assistance will be coordinated by the manager.

  • Ewa explained that local market teams must log server and DNS requests in ServiceNow and drive their own migration processes, with operations support provided by the available person rather than a dedicated resource.

ServiceNow integration

  • Ewa stated that an enhancement request will be logged to obtain API access to ServiceNow, enabling better integration and automation for database and monitoring tasks.

Splunk deployment

  • Jimmy reported that the Splunk deployment project is nearing completion, with disk monitoring redesigned to use dynamic thresholds instead of macros, and the team plans to review the new approach with all operations teams.

Disk monitoring improvements

  • Jimmy demonstrated the new disk monitoring thresholds, which use dynamic levels based on drive size instead of fixed percentages.

  • Jimmy explained that the new alert system applies default thresholds to all servers, with exceptions possible, and alerts are triggered based on specific free space levels and durations.

Disk space alert routing

  • Jimmy agreed to investigate whether incidents can be configured to route to different teams based on whether the issue is with the system drive or data drive.

  • Participants discussed sending warning-level disk space alerts to server owners and critical-level alerts directly to the operations team, with ECC monitoring after hours.

Disk space alert escalation

  • Leszek raised the question of whether X Matters should be used to notify the on-call team directly for critical disk space alerts, bypassing ECC.

Disk monitoring and alerting

  • Participants discussed that disk space incidents and alerts are only generated for production systems, while non-production systems have alerts in Zabbix but do not generate incidents.

  • Jimmy explained that exceptions to the default disk monitoring rules can be managed by tagging specific disks in the file inventory, allowing for custom detectors and alert thresholds.

  • Participants agreed that incidents should be created when monitoring agents stop responding, to prevent these issues from being overlooked, and discussed whether the incident should be triggered after a certain number of hours or days.

Monitoring and server management

  • Participants agreed that incidents should be created if a monitoring agent is unresponsive for one hour on production systems.

  • Mats and Jimmy explained that detector configurations and dashboards in Splunk are managed by code and can be restored or updated automatically, with documentation available via a shared link.

Follow-up tasks

Task Assigned to Due date Bucket
Discuss certificate replacement responsibilities and global practices with Beata and Jeff (Adrian, Lesha)
Uninstall Oracle Java from all company servers as directed by the legal team (Leschek, Adrian)
Decide whether to generate tickets for warning-level disk space alerts (80%) or only for critical-level alerts (90%) (participants)
Decide whether disk space alert tickets should be sent to the server owner or directly to the operations team (participants)
Add maintenance mode step to the server decommissioning process before powering off servers (Jimmy)