May 17, 2026, 11:40 PM
This commit is contained in:
@@ -0,0 +1,149 @@
|
||||
---
|
||||
type: 'DailyNote'
|
||||
title: October 7, 2025
|
||||
date: '2025-10-07T00:00:00.000Z'
|
||||
tags: [dba, migracja]
|
||||
---
|
||||
|
||||
### Uporządkowane notatki ze spotkania
|
||||
|
||||
#### 1) Temat przewodni
|
||||
|
||||
- Przegląd zadań infrastrukturalnych i wpływu na DBA
|
||||
|
||||
- Szacunki czasowe i odpowiedzialności zespołów (DBA, Windows, Linux, Network)
|
||||
|
||||
- Porządkowanie procesu ticketów/obserwowalności i dokumentacji pracy
|
||||
|
||||
- Planowanie rebootów (server uptime) i migracji (m.in. HU)
|
||||
|
||||
#### 2) Decyzje i ustalenia
|
||||
|
||||
- Active Directory end-of-life:
|
||||
|
||||
- Decommissioning kontrolerów domeny w Europie (ok. 91 DC) – brak wpływu na bazy danych, bo nie dekomisjonujemy domen/AD.
|
||||
|
||||
- DHCP, KMS, serwery plików, Veeam, jump stations (Windows):
|
||||
|
||||
- Brak bezpośrednich zadań dla DBA, z wyjątkiem jump stations.
|
||||
|
||||
- Jump stations (Windows):
|
||||
|
||||
- Potrzebne klienty DB: SQL Server Management Studio (SSMS), MySQL client/workbench (freeware, ale wymagane zgłoszenie/proces).
|
||||
|
||||
- Konieczna migracja zapisanych połączeń/konfiguracji (uwaga: skopiować pełny zestaw plików konfiguracyjnych, nie 2/3).
|
||||
|
||||
- Integracja z CyberArk – zachować spójność po migracji.
|
||||
|
||||
- Szacunek pracy: wpisujemy 20 godzin (kompromis – w zespole są rozbieżne potrzeby od 2h do pełnego dnia/osobę).
|
||||
|
||||
- Server uptime / kwartalne restarty:
|
||||
|
||||
- Cel: każdy serwer restartowany raz na kwartał; przygotowanie procesu identyfikacji maszyn z długim uptime i bezpiecznych restartów.
|
||||
|
||||
- DBA: obecność standby podczas okien restartowych, weryfikacja dostępności instancji po restarcie, szybkie sanity checki.
|
||||
|
||||
- Szacunek przygotowawczy: 30 godzin (osobodzień DBA/kwartał na przeglądy i koordynację; BAU/KTLO później).
|
||||
|
||||
- Migracja HU (Węgry):
|
||||
|
||||
- 16 serwerów (gł. Windows). Zakres DBA zależny od tego, czy to będzie „lift-and-shift” VM czy równoległa modernizacja (np. 2012 → 2025).
|
||||
|
||||
- Jeśli 1:1 przeniesienie – kilka godzin DBA na walidację; jeśli modernizacja – większy nakład (kontrola kompatybilności, logins, joby, ustawienia instancji).
|
||||
|
||||
- Na teraz wpisane 50 godzin (do weryfikacji po doprecyzowaniu listy baz i sposobu migracji).
|
||||
|
||||
- Projekty end-of-life baz danych:
|
||||
|
||||
- Nie ma dodatkowych pozycji poza tymi widocznymi; migracje baz trwają ciągle (co miesiąc). To powinno być uwzględniane w planie rocznym.
|
||||
|
||||
- Inne:
|
||||
|
||||
- „Disable switchboards impaired by migration to GCP” – faza discovery przez ServiceNow/DB app (czeka na doprecyzowanie/scan).
|
||||
|
||||
#### 3) Wpływ na DBA – główne punkty
|
||||
|
||||
- Jump stations:
|
||||
|
||||
- Instalacja i konfiguracja klientów (SSMS, MySQL), przeniesienie profili/połączeń, dostosowanie CyberArk.
|
||||
|
||||
- Rebooty kwartalne:
|
||||
|
||||
- Standby i post-check po restarcie; test zapytań zdrowotnych; ewentualne odtworzenie dostępu.
|
||||
|
||||
- Migracje (ciągłe, miesięczne):
|
||||
|
||||
- Realny, stały nakład pracy nieodzwierciedlony w bieżącym arkuszu – potrzeba lepszej ewidencji.
|
||||
|
||||
- HU:
|
||||
|
||||
- Do potwierdzenia sposób migracji i lista instancji, możliwy większy effort przy modernizacji.
|
||||
|
||||
#### 4) Szacunki godzin (aktualne wpisy)
|
||||
|
||||
- Jump stations (DBA): 20h
|
||||
|
||||
- Server uptime (przygotowanie procesu, nie BAU): 30h
|
||||
|
||||
- Migracja HU: 50h
|
||||
|
||||
- Inne zespoły: w arkuszu łącznie ~60h dla Network/Linux/Windows (potwierdzone jako sensowne)
|
||||
|
||||
- Uwaga: migracje comiesięczne baz – duży, stały effort, którego nie widać w sumie godzin (do odzwierciedlenia).
|
||||
|
||||
#### 5) Ryzyka i uwagi
|
||||
|
||||
- Utrata zapisanych połączeń/konfiguracji na jump stations – konieczny pełny backup plików konfiguracyjnych klientów DB.
|
||||
|
||||
- Zależności od AD (konty serwisowe dla SQL/Agent), uprawnień i dostępów – czasochłonne i często niewidoczne w taskach głównych.
|
||||
|
||||
- Niska obserwowalność po stronie DB (brak pełnej integracji ze Splunk/Zabbix/ServiceNow) powoduje niedoszacowanie prac DBA.
|
||||
|
||||
- Długie cykle ticketów DBA i prace równoległe utrudniają rzetelne raportowanie czasu.
|
||||
|
||||
#### 6) Działania następne
|
||||
|
||||
- Jump stations:
|
||||
|
||||
- Sporządzić krótką instrukcję eksportu/importu konfiguracji klientów (SSMS, MySQL Workbench/CLI).
|
||||
|
||||
- Zgłosić instalacje freeware formalnym kanałem; potwierdzić zgodność z CyberArk.
|
||||
|
||||
- Server uptime:
|
||||
|
||||
- Zbudować proces: generowanie listy serwerów z długim uptime, planowe okna, lista kontrolna DB po restarcie, przypisanie odpowiedzialnych/standby.
|
||||
|
||||
- Migracja HU:
|
||||
|
||||
- Uzyskać inwentarz baz/instancji, potwierdzić tryb migracji (lift-and-shift vs modernizacja), zaktualizować estymaty.
|
||||
|
||||
- Widoczność pracy DBA:
|
||||
|
||||
- Ustalić automatyzację generowania ticketów z alertów (preferowany Splunk; ewentualnie mostkowanie Zabbix → ServiceNow).
|
||||
|
||||
- Zdefiniować minimalne logowanie czasu przy incydentach/projektach, aby tworzyć metryki (dashboard obciążenia i typów pracy).
|
||||
|
||||
- Przegląd listy odpowiedzialności DBA i dopisanie brakujących elementów do arkusza projektów (szczególnie migracje cykliczne).
|
||||
|
||||
- Komunikacja:
|
||||
|
||||
- W piątek wrócić do tematu metryk/ticketów i potwierdzić estymaty (zwłaszcza jump stations, server uptime, HU).
|
||||
|
||||
#### 7) Otwarte pytania
|
||||
|
||||
- Czy HU będzie czystym przeniesieniem VM, czy modernizacją systemów/SQL?
|
||||
|
||||
- Jaki dokładnie zestaw klientów DB ma być standardem na jump stations (wersje, polityka aktualizacji)?
|
||||
|
||||
- Które narzędzie do automatyzacji ticketów wybieramy (Splunk vs integracja Zabbix → ServiceNow) i jaki jest plan wdrożenia?
|
||||
|
||||
- Jak formalnie włączyć comiesięczne migracje DB do planu/arkusza (pozycja BAU/KTLO czy dedykowane projekty)?
|
||||
|
||||
Jeśli chcesz, przygotuję krótką checklistę dla:
|
||||
|
||||
- eksportu/importu profili połączeń SSMS i MySQL Workbench,
|
||||
|
||||
- sanity check po restarcie serwera dla SQL Server i MySQL,
|
||||
|
||||
- szablon zadania w ServiceNow dla „quarterly reboot DB check”.
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
Meeting notes
|
||||
Incident management
|
||||
- Mariusz reported receiving multiple daily Splunk alerts about disk usage, resulting in duplicate incidents that affect statistics.
|
||||
- Mariusz fixed eight database-related disk usage cases and escalated two cases to the Windows team for disk extension.
|
||||
- Mariusz explained that the Windows team requires a five-day advance change request for disk extensions, causing delays in resolving incidents.
|
||||
- Mariusz clarified that duplicate daily incidents are being generated from Zabbix via Splunk integration, not from Splunk observability alerts.
|
||||
- Oleksandr explained that Splunk updates existing incidents instead of creating new tickets, while Zabbix integration creates multiple tickets due to its configuration.
|
||||
- Mariusz described ongoing issues with duplicate incident tickets generated daily, which negatively affect team statistics.
|
||||
System integration
|
||||
- Mats mentioned ongoing testing of ServiceNow integration, which may be related to the incident creation process.
|
||||
- Mariusz explained that some servers in the failover database cluster were never attached to Splunk due to port conflicts, despite the Splunk agent being installed.
|
||||
- Mariusz described ongoing errors when attempting to update monitoring agents or add servers to Ansible, including connection failures and outdated agents.
|
||||
- Kamil offered to mark servers missing in Splunk in Mariusz's file to help prioritize which servers to skip for now
|
||||
System monitoring gaps
|
||||
- Mariusz highlighted that several servers are not monitored by basic infrastructure templates, making it impossible to add database monitoring.
|
||||
- Mats reported sending a ticket to Lashek after discovering many Windows servers missing from monitoring, while Linux servers were fully accounted for.
|
||||
- Mats reported sending a ticket to Lashek after discovering many Windows servers missing from monitoring, while Linux servers were fully accounted for.
|
||||
System troubleshooting
|
||||
- Oleksandr suggested checking firewall settings and block lists as a possible cause for recent connection issues, but Mariusz noted the servers had worked a week ago.
|
||||
- Oleksandr explained that connection issues may be due to disabled services or local firewall settings on Windows machines, and suggested testing with the Windows team
|
||||
Follow-up tasks
|
||||
| Task | Assigned to | Due date | Bucket |
|
||||
| :-------------------------------------------------------------------------------------------------------------------------------------- | :---------- | :------- | :----- |
|
||||
| Mariusz stated they will discuss the issue of excessive incident tickets with Adrian to address the impact on team statistics (Mariusz) | | | |
|
||||
| Contact the Windows team to resolve Ansible connection issues for specific servers (Mariusz) | | | |
|
||||
@@ -0,0 +1,77 @@
|
||||
# Notatki ze spotkania (Bi-weekly DevOps)
|
||||
**Data:** 13.05.2026
|
||||
|
||||
## Cel spotkania
|
||||
Przegląd postępów w projektach i statusu operacyjnego.
|
||||
|
||||
## Decyzje
|
||||
- Dodanie dni wolnych (tzw. "Squeeze Day") dla pracowników ze Szwecji w systemie HR.
|
||||
- Zamknięcie projektu ***** i przejście do wsparcia operacyjnego.
|
||||
- Obniżenie poziomu logowania Splunk w celu zmniejszenia wolumenu logów.
|
||||
- Wykorzystanie skryptów NinjaRMM do aktualizacji plików konfiguracyjnych Splunk na serwerach Windows.
|
||||
- Rozpoczęcie procesu akceptacji dla globalnego dostępu do Remote Desktop Manager.
|
||||
|
||||
## Otwarte pytania i problemy
|
||||
- **System HR:** Wyjaśnienie kwestii szwedzkich dni wolnych i "Squeeze Days" w systemie HR.
|
||||
- **Migracja firewalli:** Zbadanie incydentów związanych z migracją firewalli oraz problemów z asymetrycznym routingiem (wycofanie zmian nie wchodzi w grę z powodu końca wsparcia i braku odnowienia licencji starych urządzeń). Celem jest zakończenie migracji wszystkich firewalli CGI (w tym OVH, AWS i Węgry) na Palo Alto, zastępując firewalle Checkpoint, do końca listopada.
|
||||
- **Licencje i patchowanie Red Hat:** Oczekiwanie na odnowienie licencji Red Hat po zakończeniu oceny ryzyka przez dostawcę. Opóźnienie wynika z wolnej odpowiedzi Red Hat na pytania. Około 10 nieprodukcyjnych serwerów Red Hat jest podatnych na nowe luki, ale ryzyko jest umiarkowane i patchowanie może poczekać w ramach 90-dniowego okna. Patchowanie planowane jest przed produkcyjnymi weekendami.
|
||||
- **Monitorowanie licencji:** Konieczność poprawy procesu monitorowania licencji i powiadomień o odnowieniach (brak alertów z systemów wewnętrznych, powiadomienia przychodziły przez DocuSign).
|
||||
|
||||
## Status projektów
|
||||
- Inżynierowie zakończyli projekt ***** (Kamil wyśle szczegółowe informacje).
|
||||
- Zespół inżynierów rozpoczyna nowe projekty, w tym globalną automatyzację wymiany certyfikatów (pod przewodnictwem Kamila) oraz prace nad agentami logowania scen.
|
||||
|
||||
## Zadania
|
||||
- [ ] **Ewa:** Dodanie dni wolnych dla wszystkich szwedzkich pracowników w systemie HR.
|
||||
|
||||
---
|
||||
## Źródło (Source)
|
||||
|
||||
Decisions
|
||||
|
||||
Add public holiday (Squeeze Day) for Swedish workers in HR system.
|
||||
Close ***** project and transition to operational support.
|
||||
Lower Splunk logging level to reduce log volume.
|
||||
Use NinjaRMM scripting to update Splunk config files on Windows servers.
|
||||
Initiate approval process for Remote Desktop Manager global access.
|
||||
Open questions
|
||||
|
||||
Clarify Swedish holidays and Squeeze Days in HR system.
|
||||
Investigate firewall migration incidents and asymmetric routing problems.
|
||||
Wait for Red Hat license renewal to complete risk assessment and patching.
|
||||
Improve license monitoring and renewal notification process.
|
||||
Agenda
|
||||
Goal: Review project progress and operational status
|
||||
|
||||
|
||||
|
||||
Ewa updated the HR system to include the public holiday for Swedish workers, addressing the confusion about the upcoming days off.
|
||||
Participants discussed the concept of "squeeze days" in Sweden, clarifying that these are days off between holidays and regular workdays, which are not common in Poland.
|
||||
Ewa updated the HR system to include the public holiday for Swedish workers and planned to notify the Swedish team in their chat.
|
||||
Project updates
|
||||
|
||||
Ewa announced that the engineering team completed the ***** project, and Kamil will send out more information to relevant people.
|
||||
Ewa explained that the engineering team is starting new projects, including automating certificate replacement globally with Kamil leading, and working on scene logging agents.
|
||||
Firewall migration
|
||||
|
||||
Participants discussed ongoing issues with undocumented firewall flows and unexpected failures during migration, noting that rollback is not an option due to end-of-life and non-renewal of licenses.
|
||||
Ewa stated that the goal is to complete migration of all CGI firewalls, including OVH, AWS, and Hungary, to Palo Alto by the end of November, replacing Checkpoint firewalls.
|
||||
Splunk logging
|
||||
|
||||
Leszek raised concerns about excessive Splunk logging obscuring problem analysis, and Jimmy suggested lowering the logging level and updating the config file across Windows servers.
|
||||
Red Hat licensing and patching
|
||||
|
||||
Adrian explained that Red Hat license expiration is preventing patching of affected servers, and renewal is pending completion of Red Hat's risk assessment, with patching planned before production weekends.
|
||||
Red Hat licensing
|
||||
|
||||
Adrian explained that the Red Hat license renewal process was delayed due to Red Hat's slow response to assessment questions, but he expects the contract to be renewed before production patching begins.
|
||||
Red Hat vulnerabilities
|
||||
|
||||
Adrian stated that there are about 10 non-production Red Hat servers affected by the new vulnerability, but the risk is moderate and patching can wait within the allowed 90-day window.
|
||||
License monitoring
|
||||
|
||||
Leszek asked about tools for monitoring license status, and Adrian clarified that notifications are typically received via DocuSign, but internal systems did not provide alerts for this renewal.
|
||||
Follow-up tasks
|
||||
|
||||
Task Assigned to Due date Bucket
|
||||
Add the public holiday for all Swedish workers in the HR system (Ewa)
|
||||
@@ -0,0 +1,86 @@
|
||||
📝 Notatka ze spotkania: Local Swiss System Piccard
|
||||
Cel spotkania: Ujednolicenie oczekiwań oraz planu działań w związku z niedawnymi incydentami wydajnościowymi w systemie Piccard i innych systemach szwajcarskich.
|
||||
|
||||
🔍 Główne tematy dyskusji
|
||||
Wpływ incydentów i odpowiedzialność: Ostatnie awarie miały bezpośredni wpływ na ciągłość biznesową w Szwajcarii (Antje). Ustalono, że konieczne jest jasne zdefiniowanie własności systemów i odpowiedzialności za rozwiązywanie problemów.
|
||||
Zmiany organizacyjne i dług techniczny: Zespoły Sundarama, Radka i Luki raportują teraz do Sunila, a zespoły infrastrukturalne do Adriana. Po odejściu Guido (który dbał o doraźne łatki) zespół musi skupić się na eliminacji przyczyn źródłowych (Root Cause) zamiast stosowania obejść.
|
||||
Zarządzanie incydentami: Należy poprawić jakość zgłoszeń (ticketów) i wykorzystywać tzw. "working bridges" (mostki) w celu szybszej i lepszej współpracy międzyzespołowej. Aby priorytetyzować prace nad przestarzałymi systemami, Antje zasugerowała mocniejsze wykorzystywanie procesów PIR (Post-Incident Review).
|
||||
Problemy z bazą danych: Trwające problemy z blokującymi się sesjami i stabilnością dotyczą bazy audytowej frameworka Spring Batch, a nie głównej bazy biznesowej Piccard (Martin, Luca). Jest to powrót problemu sprzed 3 lat.
|
||||
Propozycje rozwiązań technicznych:
|
||||
Głęboka analiza pojemności i zmian we wzorcach użycia bazy audytowej (Antje).
|
||||
Partycjonowanie danych – rozwiązanie, które w przeszłości zastosował Guido, gdy wolumen danych znacząco wzrósł (Luca).
|
||||
Analiza zależności procesów, odpytywanie bazy o uruchomione zadania i ich sortowanie, aby zidentyfikować główne blokady (Leszek).
|
||||
Aktualizacja systemu Piccard: Kompleksowy upgrade to duży projekt wymagający zmian strukturalnych w nowej bazie danych (Roberto).
|
||||
🎯 Podjęte decyzje
|
||||
Wszystkie stopniowe ulepszenia i wnioski z analiz będą na bieżąco dokumentowane w systemie śledzenia zadań.
|
||||
Zostanie zaplanowane kolejne spotkanie (follow-up) w celu weryfikacji postępów prac.
|
||||
❓ Otwarte pytania
|
||||
Dokładna przyczyna źródłowa (Root Cause) ciągłego blokowania się bazy audytowej nadal wymaga zbadania.
|
||||
Decyzja o aktualizacji (upgrade) całego systemu Piccard jest zawieszona – wymaga weryfikacji dostępnych zasobów i ustalenia priorytetów.
|
||||
📋 Zadania do wykonania (Action Items)
|
||||
[Antje, Adrian, Sundaram, Radek, Luca] – Doprecyzowanie odpowiedzialności i podziału zadań związanych z trwającymi incydentami w systemach szwajcarskich.
|
||||
[Adrian, Luca, Martin, DB team, AO team] – Ustanowienie "incident bridge" dla awarii związanych z systemem Piccard w celu sprawnej współpracy i przypisywania zadań do odpowiedzialnych zespołów.
|
||||
[Luca] – Przygotowanie listy wymaganych aktualizacji systemowych i ulepszeń dla szwajcarskich produktów (i udostępnienie jej na czacie).
|
||||
[Adrian] – Rozpoczęcie analizy przyczyny źródłowej powtarzających się problemów z wykonywaniem zadań (jobów) w bazie audytowej.
|
||||
[Paweł, Leszek, Luca] – Wypracowanie jasnego podziału na problemy leżące po stronie aplikacji vs problemy po stronie bazy danych, aby usprawnić troubleshooting.
|
||||
[Zespoły techniczne] – Zaplanowanie regularnych sesji roboczych (deep dives / burze mózgów) w celu rozwiązywania cyklicznie powracających problemów.
|
||||
|
||||
|
||||
--- source
|
||||
Decisions
|
||||
|
||||
Document incremental improvements and findings in a tracking system.
|
||||
Set up a follow-up meeting to review progress.
|
||||
Open questions
|
||||
|
||||
Root cause of audit database blocking needs further investigation.
|
||||
Feasibility of Picard system upgrade is unresolved due to priorities and resources.
|
||||
Agenda
|
||||
Goal: Align on expectations and actions regarding recent incidents in Piccard and other Swiss systems
|
||||
|
||||
|
||||
Review recent incidents for Piccard and Swiss systems (10 min)
|
||||
Discuss potential improvements in monitoring and alerting (10 min)
|
||||
Align on expectations and actions for incident/engagement process (10 min)
|
||||
Meeting notes
|
||||
Incident impact and ownership
|
||||
|
||||
Antje explained that recent incidents in Piccard and other Swiss systems have caused disruptions to the local Swiss business, and highlighted the need to clarify ownership and responsibilities moving forward.
|
||||
Roles and reporting structure
|
||||
|
||||
Antje confirmed that Sundaram, Radek, Luca, and their teams now report to Sunil, while infrastructure teams continue to report to Adrian.
|
||||
System maintenance and transition
|
||||
|
||||
Adrian stated that Guido previously played a critical role in maintaining Swiss products, often implementing temporary fixes, and emphasized the need to address root causes now that Guido is no longer with the organization.
|
||||
Luca confirmed that their team is actively working on improvements for recurring issues in Swiss products
|
||||
Antje suggested using the PIR as a vehicle to prioritize backlog items for aging Swiss systems
|
||||
Incident management
|
||||
|
||||
Antje discussed the importance of improving ticket quality and encouraged all teams to use working bridges for timely cross-team collaboration
|
||||
Luca described ongoing collaboration with Paweł, Demchuk, and Martin to address daily database hanging sessions and improve system stability
|
||||
System stability and database issues
|
||||
|
||||
Martin clarified that the recurring issues are caused by the audit database used by the Spring Batch framework, not the Picard business database.
|
||||
Martin described that recurring job execution issues in the audit database have resurfaced, similar to incidents from three years ago, and emphasized that fixing these consumes significant daily effort.
|
||||
Antje suggested conducting a deeper analysis of the audit database to investigate potential capacity issues and evolving usage patterns.
|
||||
Luca recalled that Guido previously improved performance by partitioning the database when data volume reached a certain level, and proposed reviewing if similar action is needed now.
|
||||
Leszek recommended breaking down processes, querying the database for running jobs, and sorting them by data dependencies to identify and resolve blocking issues.
|
||||
Martin described that recurring job execution issues in the audit database have resurfaced, similar to incidents from three years ago, and emphasized that fixing these consumes significant daily effort.
|
||||
Antje suggested conducting a deeper analysis of the audit database to investigate potential capacity issues and evolving usage patterns.
|
||||
System upgrades and project scope
|
||||
|
||||
Roberto confirmed that structural changes to the new database would be required, making the upgrade a significant project.
|
||||
System architecture
|
||||
|
||||
Luca clarified that the audit database is related to the Spring Batch environment and not to GDPR compliance, addressing previous confusion
|
||||
Follow-up tasks
|
||||
|
||||
Task Assigned to Due date Bucket
|
||||
Clarify ownership of tasks and responsibilities for ongoing incidents in Swiss systems (Antje, Adrian, Sundaram, Radek, Luca)
|
||||
Establish an incident bridge for Picard-related incidents to enable cross-team collaboration and assign tasks to accountable teams (Adrian, Luca, Martin, DB team, AO team)
|
||||
Compile a list of system upgrades and improvements needed for Swiss products and share in the chat (Luca)
|
||||
Adrian proposed investigating the underlying cause of recurring job execution issues in the audit database after the meeting.
|
||||
Clarify the distinction between application-level and database-level issues for more effective troubleshooting (Paweł, Leszek, Luca)
|
||||
Antje suggested that the technical team hold regular working sessions to conduct deep dives and brainstorming on recurring issues
|
||||
Compile a list of system upgrades and improvements needed for Swiss products and share in the chat (Luca)
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
|
||||
Meeting notes
|
||||
Meeting cadence
|
||||
|
||||
Ewa announced that the meeting will now be held weekly instead of bi-weekly based on feedback from the Engineering and Operations teams.
|
||||
Project management
|
||||
|
||||
Ewa explained that a new "one project model" will be implemented, consolidating all initiatives and projects into a single Jira portfolio project called ONE.
|
||||
Project updates
|
||||
|
||||
Ewa provided updates on ongoing projects, including completion of the CyberArk project, upcoming SIM logging implementation in Europe, and plans to automate certificate installations with the Venify team.
|
||||
CyberArk project status
|
||||
|
||||
Ewa clarified that the CyberArk project is considered complete from the EU ONE perspective, with implementation finished for infrastructure and database engineering, but ongoing onboarding and developer support now fall under operations and security teams.
|
||||
CyberArk technical limitations
|
||||
|
||||
Oleksandr Lozinskyi explained that CyberArk implementation is complete for Windows environments, but several limitations remain for Linux, such as inability to rotate SSH keys and lack of support for port forwarding, requiring manual intervention for key replacement.
|
||||
CyberArk operational transition
|
||||
|
||||
Ewa stated that future policy decisions and developer onboarding for CyberArk will be managed by the security team, with requests handled through ticketing rather than ongoing project involvement.
|
||||
Ewa stated that future policy decisions and developer onboarding for CyberArk are now the responsibility of the security team, and any new requirements will be handled as separate projects through ticketing.
|
||||
CyberArk access policies
|
||||
|
||||
Jimmy clarified that all users are now required to use CyberArk for initial connections to jump stations, with approval for this process in place.
|
||||
Oleksandr explained that general access remains possible, but root permissions must be obtained through CyberArk, resulting in dual management of keys for both CyberArk and general user accounts.
|
||||
Jimmy clarified that all users must use CyberArk for initial connections to jump stations, but in emergency incidents, a break glass account may be used with approval from Adrian or another incident bridge manager.
|
||||
Splunk alert management
|
||||
|
||||
Jimmy explained that disk alert incidents in Splunk can be grouped into a single incident by standardizing the alert subject naming, and new alerts will add comments to the existing incident.
|
||||
Jimmy stated that while incidents cannot be assigned to specific teams in Splunk, X-Matters callouts can be configured to notify the appropriate team based on the type of disk alert.
|
||||
Jimmy clarified that incidents in Splunk cannot be directly assigned to specific teams, but X-Matters callouts can be configured to notify the appropriate team based on the type of disk alert.
|
||||
Incident response and monitoring
|
||||
|
||||
Adrian described a recent incident in Switzerland where database disk alerts were not acted upon for three weeks, resulting in disk space exhaustion and questions about monitoring redundancy.
|
||||
Adrian stated that leadership in the US recommended database disk alerts should always go to the database team, as development teams often cannot resolve disk space issues.
|
||||
Incident management and accountability
|
||||
|
||||
Adrian raised concerns about accountability for disk space incidents, emphasizing that application teams should be responsible for managing their own servers and infrastructure, not the infra team.
|
||||
Ewa explained that the global process assigns responsibility for disk alerts and server maintenance to application teams, and noted that this approach will be reinforced as teams report to Sunil.
|
||||
Disk space accountability
|
||||
|
||||
Adrian stated that DRN owners should be held accountable for disk space issues, not the infrastructure team, and suggested that clear communication is needed to reinforce this responsibility.
|
||||
Disk alert management
|
||||
|
||||
Oleksandr reported that 56 tickets related to disk errors remain unattended by the application team, highlighting ongoing issues with disk alert management on Linux systems.
|
||||
Communication to application teams
|
||||
|
||||
Ewa confirmed that application teams were previously informed about disk alert processes and responsibilities through meetings, emails, and Yammer posts, and that reminders have been sent recently.
|
||||
Incident assignment process
|
||||
|
||||
Ewa stated that Splunk-generated incidents are assigned to application teams for initial investigation and validation, with further escalation if needed, and that this process aligns with the European global process.
|
||||
Alert escalation improvements
|
||||
|
||||
Jimmy proposed adding additional alert levels and X-Matters callouts to notify both operations and application teams when disk space issues approach critical levels.
|
||||
Application deployment requests
|
||||
|
||||
Ewa shared that the application team requested code implementation, and participants discussed that AO should handle deployment, not Operations or Engineering
|
||||
Jimmy explained that infrastructure teams should build and maintain their own solutions for requests, rather than running scripts provided by development teams
|
||||
Follow-up tasks
|
||||
|
||||
Task Assigned to Due date Bucket
|
||||
Create documentation for processing CyberArk requests and ensure it covers details for all ADs and the Splunk process (Konrad, Leszek)
|
||||
Adrian suggested that communication should be sent to clarify that DRN owners are accountable for disk space issues, not the infrastructure team.
|
||||
Review the application team's code deployment request and determine if AO or Linux team should handle deployment (Leszek)
|
||||
|
||||
Reference in New Issue
Block a user