71 lines
5.4 KiB
Markdown
71 lines
5.4 KiB
Markdown
---
|
|
type: Meeting
|
|
Status: Archived
|
|
_archived: true
|
|
---
|
|
|
|
Meeting notes
|
|
Meeting cadence
|
|
|
|
Ewa announced that the meeting will now be held weekly instead of bi-weekly based on feedback from the Engineering and Operations teams.
|
|
Project management
|
|
|
|
Ewa explained that a new "one project model" will be implemented, consolidating all initiatives and projects into a single Jira portfolio project called ONE.
|
|
Project updates
|
|
|
|
Ewa provided updates on ongoing projects, including completion of the CyberArk project, upcoming SIM logging implementation in Europe, and plans to automate certificate installations with the Venify team.
|
|
CyberArk project status
|
|
|
|
Ewa clarified that the CyberArk project is considered complete from the EU ONE perspective, with implementation finished for infrastructure and database engineering, but ongoing onboarding and developer support now fall under operations and security teams.
|
|
CyberArk technical limitations
|
|
|
|
Oleksandr Lozinskyi explained that CyberArk implementation is complete for Windows environments, but several limitations remain for Linux, such as inability to rotate SSH keys and lack of support for port forwarding, requiring manual intervention for key replacement.
|
|
CyberArk operational transition
|
|
|
|
Ewa stated that future policy decisions and developer onboarding for CyberArk will be managed by the security team, with requests handled through ticketing rather than ongoing project involvement.
|
|
Ewa stated that future policy decisions and developer onboarding for CyberArk are now the responsibility of the security team, and any new requirements will be handled as separate projects through ticketing.
|
|
CyberArk access policies
|
|
|
|
Jimmy clarified that all users are now required to use CyberArk for initial connections to jump stations, with approval for this process in place.
|
|
Oleksandr explained that general access remains possible, but root permissions must be obtained through CyberArk, resulting in dual management of keys for both CyberArk and general user accounts.
|
|
Jimmy clarified that all users must use CyberArk for initial connections to jump stations, but in emergency incidents, a break glass account may be used with approval from Adrian or another incident bridge manager.
|
|
Splunk alert management
|
|
|
|
Jimmy explained that disk alert incidents in Splunk can be grouped into a single incident by standardizing the alert subject naming, and new alerts will add comments to the existing incident.
|
|
Jimmy stated that while incidents cannot be assigned to specific teams in Splunk, X-Matters callouts can be configured to notify the appropriate team based on the type of disk alert.
|
|
Jimmy clarified that incidents in Splunk cannot be directly assigned to specific teams, but X-Matters callouts can be configured to notify the appropriate team based on the type of disk alert.
|
|
Incident response and monitoring
|
|
|
|
Adrian described a recent incident in Switzerland where database disk alerts were not acted upon for three weeks, resulting in disk space exhaustion and questions about monitoring redundancy.
|
|
Adrian stated that leadership in the US recommended database disk alerts should always go to the database team, as development teams often cannot resolve disk space issues.
|
|
Incident management and accountability
|
|
|
|
Adrian raised concerns about accountability for disk space incidents, emphasizing that application teams should be responsible for managing their own servers and infrastructure, not the infra team.
|
|
Ewa explained that the global process assigns responsibility for disk alerts and server maintenance to application teams, and noted that this approach will be reinforced as teams report to Sunil.
|
|
Disk space accountability
|
|
|
|
Adrian stated that DRN owners should be held accountable for disk space issues, not the infrastructure team, and suggested that clear communication is needed to reinforce this responsibility.
|
|
Disk alert management
|
|
|
|
Oleksandr reported that 56 tickets related to disk errors remain unattended by the application team, highlighting ongoing issues with disk alert management on Linux systems.
|
|
Communication to application teams
|
|
|
|
Ewa confirmed that application teams were previously informed about disk alert processes and responsibilities through meetings, emails, and Yammer posts, and that reminders have been sent recently.
|
|
Incident assignment process
|
|
|
|
Ewa stated that Splunk-generated incidents are assigned to application teams for initial investigation and validation, with further escalation if needed, and that this process aligns with the European global process.
|
|
Alert escalation improvements
|
|
|
|
Jimmy proposed adding additional alert levels and X-Matters callouts to notify both operations and application teams when disk space issues approach critical levels.
|
|
Application deployment requests
|
|
|
|
Ewa shared that the application team requested code implementation, and participants discussed that AO should handle deployment, not Operations or Engineering
|
|
Jimmy explained that infrastructure teams should build and maintain their own solutions for requests, rather than running scripts provided by development teams
|
|
Follow-up tasks
|
|
|
|
Task Assigned to Due date Bucket
|
|
Create documentation for processing CyberArk requests and ensure it covers details for all ADs and the Splunk process (Konrad, Leszek)
|
|
Adrian suggested that communication should be sent to clarify that DRN owners are accountable for disk space issues, not the infrastructure team.
|
|
Review the application team's code deployment request and determine if AO or Linux team should handle deployment (Leszek)
|
|
|