May 17, 2026, 11:40 PM
This commit is contained in:
@@ -0,0 +1,227 @@
|
||||
---
|
||||
title: 2026-04-22 środa
|
||||
type: DailyNote
|
||||
tags: [daily, 2026]
|
||||
---
|
||||
|
||||
|
||||
MIM Bridge Details
|
||||
- INC #: INC0798812
|
||||
- Short Description: Piccard Database - Degraded Performance - Switzerland
|
||||
- Priority: P3
|
||||
- Major Incident Accepted: Yes
|
||||
- Impact Start Time: TBD
|
||||
- Incident Resolution Time:
|
||||
- SNOW SOW Link: https://dnbgatewayprod.service-now.com/now/sow/record/incident/b845b8db2f908750be82f34fafa4e36c/params/selected-tab-index/3/selected-tab/id%3Dcllanqdxq01av3b7srbf5i3t3%5Eroute%3Dmim-details
|
||||
ECC2API named the meeting ECC P3 - INC0798812 - Piccard Database - Degraded Performance - Switzerland.
|
||||
|
||||
9:23 AM Meeting started
|
||||
|
||||
Facilitator has been turned on and is taking notes. How can you help us?
|
||||
|
||||
MIM Bridge Details
|
||||
- INC #: INC0798812
|
||||
- Short Description: Piccard Database - Degraded Performance - Switzerland
|
||||
- Priority: P3
|
||||
- Major Incident Accepted: Yes
|
||||
- Impact Start Time: 4/21/2026 19:35 EST
|
||||
- Incident Resolution Time: 4/22/2026 10:05 AM CET / 4:05 AM EST
|
||||
|
||||
- SNOW SOW Link: https://dnbgatewayprod.service-now.com/now/sow/record/incident/b845b8db2f908750be82f34fafa4e36c/params/selected-tab-index/3/selected-tab/id%3Dcllanqdxq01av3b7srbf5i3t3%5Eroute%3Dmim-details
|
||||
Prod
|
||||
|
||||
08:50 disk space was added
|
||||
Tentative issue start time 4/21/2026 19:35 EST
|
||||
first warning SLANK: 0:45, issue started: 1:30, slank notification: 1:35, escalated to DBA team: 6:23, resolved: 8:50-8:55
|
||||
I have another meeting and am leaving this one. Unfortunately, I can't offer much support here.
|
||||
Best regards, Géraldine
|
||||
|
||||
Sure, thank you Ricci-Rubera, Geraldine
|
||||
bisclxpbisjb01.bisnodech.local
|
||||
|
||||
1. Issue overview:
|
||||
Primary issue: Piccard application database experienced disk space exhaustion, leading to database locking sessions and severe performance degradation.
|
||||
The database ran out of disk space, preventing normal operations and causing application slowdowns and blockages.
|
||||
Although additional disk space was subsequently added, residual locking sessions in the database continued to impact performance.
|
||||
These locking sessions were identified as Hibernate sessions not closing correctly at the application layer, requiring application-level intervention.
|
||||
2. Timeline (as discussed / confirmed)
|
||||
~00:45 CET – First disk space warning received (20% remaining).
|
||||
~01:30–01:35 CET – Disk became full; Splunk alert at ~01:35 CET confirmed 100% disk usage.
|
||||
~06:23 CET – Issue escalated to DBA team.
|
||||
08:50 CET – Additional disk space (~50 GB) added to the database (confirmed multiple times by Paweł/Adrian, also reflected in meeting chat).
|
||||
Post disk extension: Database became accessible again, but locking sessions persisted, causing degraded performance.
|
||||
(Exact issue start time for business impact aligned tentatively to 21 Apr 2026 ~19:35 EST, per meeting chat.)
|
||||
3. Customer / Business Impact
|
||||
Switzerland business users heavily impacted; Piccard is widely used by most Swiss business units.
|
||||
Key impacts reported:
|
||||
Users could log in, but:
|
||||
Navigation extremely slow.
|
||||
Every click took significant time.
|
||||
Saving changes mostly hung indefinitely (loading cursor, no completion).
|
||||
Debt reports and Trade Register notifications could not be loaded earlier in the day.
|
||||
Import / shop jobs were not functioning during the incident window.
|
||||
User confirmation (Daniela Pispico):
|
||||
Application technically accessible.
|
||||
Performance “very, very slow”.
|
||||
Not usable for productive work.
|
||||
Business users confirmed they were still blocked operationally despite disk extension.
|
||||
4. Current Technical State (at time of discussion)
|
||||
✅ Disk space issue mitigated temporarily by adding capacity.
|
||||
❌ Database still affected by:
|
||||
Blocking sessions.
|
||||
Long‑running Hibernate sessions not closing.
|
||||
❌ Application performance remains degraded.
|
||||
⚠️ Working copy replication not running and intentionally deferred to avoid further business impact during working hours.
|
||||
Risk identified:
|
||||
Disk space growth has been persistently high (>80%) for some time, not a one‑off condition.
|
||||
Sudden growth may also be linked to logs, dumps, or background processes, but no final RCA yet.
|
||||
5. Key Technical Discussions / Decisions
|
||||
Restart of Tomcat process identified as the next critical remediation step:
|
||||
Purpose: Clear blocking sessions and stuck Hibernate connections.
|
||||
Clarified explicitly that this is a Tomcat process restart, not a full server reboot.
|
||||
Confirmed:
|
||||
Restart should not negatively impact other applications.
|
||||
Not all Piccard-related services are isolated, but risk is acceptable and understood.
|
||||
Linux / AO Engineering involvement required to:
|
||||
Prepare change.
|
||||
Execute controlled Tomcat restart.
|
||||
Database replication:
|
||||
Acknowledged as necessary.
|
||||
Agreed to postpone replication until end of business day to avoid further load and risk during peak usage.
|
||||
Disk space must be closely monitored before replication is resumed.
|
||||
6. Actions Being Taken (with ownership)
|
||||
Immediate / In Progress
|
||||
Add disk space to stabilize database
|
||||
✅ Completed at ~08:50 CET.
|
||||
Owner: Paweł (DBA)
|
||||
Prepare and execute Tomcat process restart
|
||||
To clear blocking and Hibernate sessions.
|
||||
Owners:
|
||||
Martin Woerner – identify affected Tomcat instance(s).
|
||||
Oleksandr Lozinskyi / AO Linux team – prepare change and restart.
|
||||
Status: Pending / being prepared during the meeting.
|
||||
Validate application behaviour post‑restart
|
||||
Confirm:
|
||||
Navigation speed.
|
||||
Ability to save changes.
|
||||
Job execution.
|
||||
Owners:
|
||||
Business users (Daniela / Alessandro / DNO team).
|
||||
IT Dev Switzerland.
|
||||
Near‑Term / Planned
|
||||
Resume database replication
|
||||
Planned post business hours (around end of day).
|
||||
Condition: Ensure sufficient disk space and stability.
|
||||
Owners: DBA + IT Dev Switzerland
|
||||
Monitor disk usage closely
|
||||
During the rest of the day and overnight while replication/jobs run.
|
||||
Owners: DBA / Engineering
|
||||
Follow‑up / Preventive
|
||||
Disk space monitoring and alerting review
|
||||
Discussion raised that:
|
||||
Disk‑full incidents are preventable and unacceptable operationally.
|
||||
Proposal to raise incidents or bridges when disk reaches 98–99% usage.
|
||||
To be taken offline as a post‑incident / peer review action.
|
||||
Owners: Engineering / DBA leadership (to be defined)
|
||||
7. Open Points / Risks
|
||||
Root cause of:
|
||||
Rapid disk consumption.
|
||||
Long‑standing high utilization trend.
|
||||
Need confirmation whether:
|
||||
Logs, dumps, backups, or specific jobs caused the sudden growth.
|
||||
Risk of disk filling up again once replication and jobs restart if not addressed structurally.
|
||||
8. Overall Status (at meeting point)
|
||||
Incident not yet resolved.
|
||||
Service is in a degraded, partially usable state.
|
||||
Restoration path agreed:
|
||||
Clear blocking via Tomcat restart.
|
||||
Validate performance with business.
|
||||
Resume replication safely.
|
||||
ECC bridge to remain active until performance and usability are confirmed restored.
|
||||
|
||||
CHG0136776
|
||||
|
||||
I’ll add this incident to the master problem we have open already
|
||||
[olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl status tomcat.service
|
||||
Wed Apr 22 10:04:21 CEST 2026
|
||||
● tomcat.service - Apache Tomcat Web Application
|
||||
Loaded: loaded (/etc/systemd/system/tomcat.service; enabled; vendor preset: disabled)
|
||||
Active: active (running) since Tue 2026-04-21 15:13:25 CEST; 18h ago
|
||||
Process: 11443 ExecStop=/usr/share/tomcat/bin/shutdown.sh (code=exited, status=0/SUCCESS)
|
||||
Process: 11526 ExecStart=/usr/share/tomcat/bin/startup.sh (code=exited, status=0/SUCCESS)
|
||||
Main PID: 11539 (java)
|
||||
CGroup: /system.slice/tomcat.service
|
||||
└─11539 /usr/bin/java -Djava.util.logging.config.file=/usr/share/tomcat/conf/logging.properties -Djava.util.logging.manager=org.apache....
|
||||
Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Starting Apache Tomcat Web Application...
|
||||
Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application.
|
||||
[olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl restart tomcat.service
|
||||
Wed Apr 22 10:04:55 CEST 2026
|
||||
[olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl status tomcat.service
|
||||
Wed Apr 22 10:05:04 CEST 2026
|
||||
● tomcat.service - Apache Tomcat Web Application
|
||||
Loaded: loaded (/etc/systemd/system/tomcat.service; enabled; vendor preset: disabled)
|
||||
Active: active (running) since Wed 2026-04-22 10:05:01 CEST; 3s ago
|
||||
Process: 18088 ExecStop=/usr/share/tomcat/bin/shutdown.sh (code=exited, status=0/SUCCESS)
|
||||
Process: 18201 ExecStart=/usr/share/tomcat/bin/startup.sh (code=exited, status=0/SUCCESS)
|
||||
Main PID: 18218 (java)
|
||||
CGroup: /system.slice/tomcat.service
|
||||
└─18218 /usr/bin/java -Djava.util.logging.config.file=/usr/share/tomcat/conf/logging.properties -Djava.util.logging.manager=org.apache.juli.Class...
|
||||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local systemd[1]: Starting Apache Tomcat Web Application...
|
||||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local startup.sh[18201]: Existing PID file found during start.
|
||||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local startup.sh[18201]: Removing/clearing stale PID file.
|
||||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application.
|
||||
[olelozloc@bisclxpbisjb01 ~]$
|
||||
|
||||
Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application.
|
||||
[olelozloc@bisclxpbisjb01 ~]$
|
||||
Tomcat restart completed
|
||||
no locks on database side
|
||||
Tentative restoration time 4/22/2026 10:05 AM CET / 4:05 AM EST
|
||||
Audit database: normal state - no locks, 25 connection from 'Collecting_Services'
|
||||
Audit database: RED state: 6-7 locks, ~70 connection from 'Collecting_Services'
|
||||
Yesterday we had restart of tomcat service on same server but related to another incident INC0798444, used change CHG0136652
|
||||
|
||||
|
||||
|
||||
|
||||
INC0798320 -> Piccard Score is not working; we cannot perform the outsourcing (the same as yesterday). Please fix this as soon as possible and let us know. Thank you. related issue INC0798444
|
||||
|
||||
INC0797772 INC0797647 INC0798267 INC0797415 INC0797747 INC0797757 INC0797445 INC0797682 INC0797490- > (Splunk Observability) Disk = 100% [H:\] chsqlbkstoprw01.bisnodech.local - system backupowy się zapchał i przenoszę dane do S3 na OVH
|
||||
|
||||
|
||||
|
||||
Wczorajsze INC0798360 INC0798393-> (Splunk Observability) Disk > 90% [I:\] CHPICBUDB0NPW02 przeniosłem dane
|
||||
|
||||
|
||||
INC0795723 INC0797683 INC0797760 INC0797672 INC0797760 INC0797696 INC0797768 INC0796110 -> (Splunk Observability) Disk > 90% [M:] CHPICBUDB0NPW01 - nie było konieczności powiększyłem dysk o 50 ponieważ nie mogłem odtworzyć bazy piccard
|
||||
|
||||
INC0798188 INC0798497 INC0797671 INC0798216-> (Splunk Observability) Disk > 90% [P:] CHPICBUDB0PRW01 - wykonałem shrink logu dla bazy Sync na piccard
|
||||
|
||||
|
||||
INC0798748 INC0798753- (Splunk Observability) Disk > 80% [J:] CHPICBUDB0PRW01 - nie potrzeba nic robić bo to tempdb
|
||||
|
||||
|
||||
INC0797669 -> (Splunk Observability) Disk = 100% [P:] CHPICBUDB0PRW01 - shrink databaze
|
||||
|
||||
|
||||
INC0798753 - > (Splunk Observability) Disk = 100% [J:] CHPICBUDB0PRW01- disk resized but waiting for enable replication
|
||||
|
||||
SCTASK0724812 -> Enable Sync and feed job
|
||||
|
||||
|
||||
Today tasks
|
||||
SCTASK0722293 -> Fulfillment task for Request Database Support - nie mam dostępu do bazy danych i musiałem odtworzyć z ostatniego backupu jaki mam, ale
|
||||
|
||||
|
||||
16370274
|
||||
16370274
|
||||
11516947
|
||||
10124205
|
||||
7985051
|
||||
6556451
|
||||
4771397
|
||||
3613868
|
||||
2977527
|
||||
2387681
|
||||
1543699
|
||||
862179
|
||||
Reference in New Issue
Block a user