227 lines
11 KiB
Markdown
227 lines
11 KiB
Markdown
---
|
||
title: 2026-04-22 środa
|
||
type: DailyNote
|
||
tags: [daily, 2026]
|
||
---
|
||
|
||
|
||
MIM Bridge Details
|
||
- INC #: INC0798812
|
||
- Short Description: Piccard Database - Degraded Performance - Switzerland
|
||
- Priority: P3
|
||
- Major Incident Accepted: Yes
|
||
- Impact Start Time: TBD
|
||
- Incident Resolution Time:
|
||
- SNOW SOW Link: https://dnbgatewayprod.service-now.com/now/sow/record/incident/b845b8db2f908750be82f34fafa4e36c/params/selected-tab-index/3/selected-tab/id%3Dcllanqdxq01av3b7srbf5i3t3%5Eroute%3Dmim-details
|
||
ECC2API named the meeting ECC P3 - INC0798812 - Piccard Database - Degraded Performance - Switzerland.
|
||
|
||
9:23 AM Meeting started
|
||
|
||
Facilitator has been turned on and is taking notes. How can you help us?
|
||
|
||
MIM Bridge Details
|
||
- INC #: INC0798812
|
||
- Short Description: Piccard Database - Degraded Performance - Switzerland
|
||
- Priority: P3
|
||
- Major Incident Accepted: Yes
|
||
- Impact Start Time: 4/21/2026 19:35 EST
|
||
- Incident Resolution Time: 4/22/2026 10:05 AM CET / 4:05 AM EST
|
||
|
||
- SNOW SOW Link: https://dnbgatewayprod.service-now.com/now/sow/record/incident/b845b8db2f908750be82f34fafa4e36c/params/selected-tab-index/3/selected-tab/id%3Dcllanqdxq01av3b7srbf5i3t3%5Eroute%3Dmim-details
|
||
Prod
|
||
|
||
08:50 disk space was added
|
||
Tentative issue start time 4/21/2026 19:35 EST
|
||
first warning SLANK: 0:45, issue started: 1:30, slank notification: 1:35, escalated to DBA team: 6:23, resolved: 8:50-8:55
|
||
I have another meeting and am leaving this one. Unfortunately, I can't offer much support here.
|
||
Best regards, Géraldine
|
||
|
||
Sure, thank you Ricci-Rubera, Geraldine
|
||
bisclxpbisjb01.bisnodech.local
|
||
|
||
1. Issue overview:
|
||
Primary issue: Piccard application database experienced disk space exhaustion, leading to database locking sessions and severe performance degradation.
|
||
The database ran out of disk space, preventing normal operations and causing application slowdowns and blockages.
|
||
Although additional disk space was subsequently added, residual locking sessions in the database continued to impact performance.
|
||
These locking sessions were identified as Hibernate sessions not closing correctly at the application layer, requiring application-level intervention.
|
||
2. Timeline (as discussed / confirmed)
|
||
~00:45 CET – First disk space warning received (20% remaining).
|
||
~01:30–01:35 CET – Disk became full; Splunk alert at ~01:35 CET confirmed 100% disk usage.
|
||
~06:23 CET – Issue escalated to DBA team.
|
||
08:50 CET – Additional disk space (~50 GB) added to the database (confirmed multiple times by Paweł/Adrian, also reflected in meeting chat).
|
||
Post disk extension: Database became accessible again, but locking sessions persisted, causing degraded performance.
|
||
(Exact issue start time for business impact aligned tentatively to 21 Apr 2026 ~19:35 EST, per meeting chat.)
|
||
3. Customer / Business Impact
|
||
Switzerland business users heavily impacted; Piccard is widely used by most Swiss business units.
|
||
Key impacts reported:
|
||
Users could log in, but:
|
||
Navigation extremely slow.
|
||
Every click took significant time.
|
||
Saving changes mostly hung indefinitely (loading cursor, no completion).
|
||
Debt reports and Trade Register notifications could not be loaded earlier in the day.
|
||
Import / shop jobs were not functioning during the incident window.
|
||
User confirmation (Daniela Pispico):
|
||
Application technically accessible.
|
||
Performance “very, very slow”.
|
||
Not usable for productive work.
|
||
Business users confirmed they were still blocked operationally despite disk extension.
|
||
4. Current Technical State (at time of discussion)
|
||
✅ Disk space issue mitigated temporarily by adding capacity.
|
||
❌ Database still affected by:
|
||
Blocking sessions.
|
||
Long‑running Hibernate sessions not closing.
|
||
❌ Application performance remains degraded.
|
||
⚠️ Working copy replication not running and intentionally deferred to avoid further business impact during working hours.
|
||
Risk identified:
|
||
Disk space growth has been persistently high (>80%) for some time, not a one‑off condition.
|
||
Sudden growth may also be linked to logs, dumps, or background processes, but no final RCA yet.
|
||
5. Key Technical Discussions / Decisions
|
||
Restart of Tomcat process identified as the next critical remediation step:
|
||
Purpose: Clear blocking sessions and stuck Hibernate connections.
|
||
Clarified explicitly that this is a Tomcat process restart, not a full server reboot.
|
||
Confirmed:
|
||
Restart should not negatively impact other applications.
|
||
Not all Piccard-related services are isolated, but risk is acceptable and understood.
|
||
Linux / AO Engineering involvement required to:
|
||
Prepare change.
|
||
Execute controlled Tomcat restart.
|
||
Database replication:
|
||
Acknowledged as necessary.
|
||
Agreed to postpone replication until end of business day to avoid further load and risk during peak usage.
|
||
Disk space must be closely monitored before replication is resumed.
|
||
6. Actions Being Taken (with ownership)
|
||
Immediate / In Progress
|
||
Add disk space to stabilize database
|
||
✅ Completed at ~08:50 CET.
|
||
Owner: Paweł (DBA)
|
||
Prepare and execute Tomcat process restart
|
||
To clear blocking and Hibernate sessions.
|
||
Owners:
|
||
Martin Woerner – identify affected Tomcat instance(s).
|
||
Oleksandr Lozinskyi / AO Linux team – prepare change and restart.
|
||
Status: Pending / being prepared during the meeting.
|
||
Validate application behaviour post‑restart
|
||
Confirm:
|
||
Navigation speed.
|
||
Ability to save changes.
|
||
Job execution.
|
||
Owners:
|
||
Business users (Daniela / Alessandro / DNO team).
|
||
IT Dev Switzerland.
|
||
Near‑Term / Planned
|
||
Resume database replication
|
||
Planned post business hours (around end of day).
|
||
Condition: Ensure sufficient disk space and stability.
|
||
Owners: DBA + IT Dev Switzerland
|
||
Monitor disk usage closely
|
||
During the rest of the day and overnight while replication/jobs run.
|
||
Owners: DBA / Engineering
|
||
Follow‑up / Preventive
|
||
Disk space monitoring and alerting review
|
||
Discussion raised that:
|
||
Disk‑full incidents are preventable and unacceptable operationally.
|
||
Proposal to raise incidents or bridges when disk reaches 98–99% usage.
|
||
To be taken offline as a post‑incident / peer review action.
|
||
Owners: Engineering / DBA leadership (to be defined)
|
||
7. Open Points / Risks
|
||
Root cause of:
|
||
Rapid disk consumption.
|
||
Long‑standing high utilization trend.
|
||
Need confirmation whether:
|
||
Logs, dumps, backups, or specific jobs caused the sudden growth.
|
||
Risk of disk filling up again once replication and jobs restart if not addressed structurally.
|
||
8. Overall Status (at meeting point)
|
||
Incident not yet resolved.
|
||
Service is in a degraded, partially usable state.
|
||
Restoration path agreed:
|
||
Clear blocking via Tomcat restart.
|
||
Validate performance with business.
|
||
Resume replication safely.
|
||
ECC bridge to remain active until performance and usability are confirmed restored.
|
||
|
||
CHG0136776
|
||
|
||
I’ll add this incident to the master problem we have open already
|
||
[olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl status tomcat.service
|
||
Wed Apr 22 10:04:21 CEST 2026
|
||
● tomcat.service - Apache Tomcat Web Application
|
||
Loaded: loaded (/etc/systemd/system/tomcat.service; enabled; vendor preset: disabled)
|
||
Active: active (running) since Tue 2026-04-21 15:13:25 CEST; 18h ago
|
||
Process: 11443 ExecStop=/usr/share/tomcat/bin/shutdown.sh (code=exited, status=0/SUCCESS)
|
||
Process: 11526 ExecStart=/usr/share/tomcat/bin/startup.sh (code=exited, status=0/SUCCESS)
|
||
Main PID: 11539 (java)
|
||
CGroup: /system.slice/tomcat.service
|
||
└─11539 /usr/bin/java -Djava.util.logging.config.file=/usr/share/tomcat/conf/logging.properties -Djava.util.logging.manager=org.apache....
|
||
Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Starting Apache Tomcat Web Application...
|
||
Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application.
|
||
[olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl restart tomcat.service
|
||
Wed Apr 22 10:04:55 CEST 2026
|
||
[olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl status tomcat.service
|
||
Wed Apr 22 10:05:04 CEST 2026
|
||
● tomcat.service - Apache Tomcat Web Application
|
||
Loaded: loaded (/etc/systemd/system/tomcat.service; enabled; vendor preset: disabled)
|
||
Active: active (running) since Wed 2026-04-22 10:05:01 CEST; 3s ago
|
||
Process: 18088 ExecStop=/usr/share/tomcat/bin/shutdown.sh (code=exited, status=0/SUCCESS)
|
||
Process: 18201 ExecStart=/usr/share/tomcat/bin/startup.sh (code=exited, status=0/SUCCESS)
|
||
Main PID: 18218 (java)
|
||
CGroup: /system.slice/tomcat.service
|
||
└─18218 /usr/bin/java -Djava.util.logging.config.file=/usr/share/tomcat/conf/logging.properties -Djava.util.logging.manager=org.apache.juli.Class...
|
||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local systemd[1]: Starting Apache Tomcat Web Application...
|
||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local startup.sh[18201]: Existing PID file found during start.
|
||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local startup.sh[18201]: Removing/clearing stale PID file.
|
||
Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application.
|
||
[olelozloc@bisclxpbisjb01 ~]$
|
||
|
||
Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application.
|
||
[olelozloc@bisclxpbisjb01 ~]$
|
||
Tomcat restart completed
|
||
no locks on database side
|
||
Tentative restoration time 4/22/2026 10:05 AM CET / 4:05 AM EST
|
||
Audit database: normal state - no locks, 25 connection from 'Collecting_Services'
|
||
Audit database: RED state: 6-7 locks, ~70 connection from 'Collecting_Services'
|
||
Yesterday we had restart of tomcat service on same server but related to another incident INC0798444, used change CHG0136652
|
||
|
||
|
||
|
||
|
||
INC0798320 -> Piccard Score is not working; we cannot perform the outsourcing (the same as yesterday). Please fix this as soon as possible and let us know. Thank you. related issue INC0798444
|
||
|
||
INC0797772 INC0797647 INC0798267 INC0797415 INC0797747 INC0797757 INC0797445 INC0797682 INC0797490- > (Splunk Observability) Disk = 100% [H:\] chsqlbkstoprw01.bisnodech.local - system backupowy się zapchał i przenoszę dane do S3 na OVH
|
||
|
||
|
||
|
||
Wczorajsze INC0798360 INC0798393-> (Splunk Observability) Disk > 90% [I:\] CHPICBUDB0NPW02 przeniosłem dane
|
||
|
||
|
||
INC0795723 INC0797683 INC0797760 INC0797672 INC0797760 INC0797696 INC0797768 INC0796110 -> (Splunk Observability) Disk > 90% [M:] CHPICBUDB0NPW01 - nie było konieczności powiększyłem dysk o 50 ponieważ nie mogłem odtworzyć bazy piccard
|
||
|
||
INC0798188 INC0798497 INC0797671 INC0798216-> (Splunk Observability) Disk > 90% [P:] CHPICBUDB0PRW01 - wykonałem shrink logu dla bazy Sync na piccard
|
||
|
||
|
||
INC0798748 INC0798753- (Splunk Observability) Disk > 80% [J:] CHPICBUDB0PRW01 - nie potrzeba nic robić bo to tempdb
|
||
|
||
|
||
INC0797669 -> (Splunk Observability) Disk = 100% [P:] CHPICBUDB0PRW01 - shrink databaze
|
||
|
||
|
||
INC0798753 - > (Splunk Observability) Disk = 100% [J:] CHPICBUDB0PRW01- disk resized but waiting for enable replication
|
||
|
||
SCTASK0724812 -> Enable Sync and feed job
|
||
|
||
|
||
Today tasks
|
||
SCTASK0722293 -> Fulfillment task for Request Database Support - nie mam dostępu do bazy danych i musiałem odtworzyć z ostatniego backupu jaki mam, ale
|
||
|
||
|
||
16370274
|
||
16370274
|
||
11516947
|
||
10124205
|
||
7985051
|
||
6556451
|
||
4771397
|
||
3613868
|
||
2977527
|
||
2387681
|
||
1543699
|
||
862179 |