Files
2026-05-18 06:40:19 +00:00

11 KiB
Raw Permalink Blame History

title, type, tags
title type tags
2026-04-22 środa DailyNote
daily
2026

MIM Bridge Details

9:23 AMMeeting started

Facilitator has been turned on and is taking notes. How can you help us?

MIM Bridge Details

08:50 disk space was added Tentative issue start time 4/21/2026 19:35 EST first warning SLANK: 0:45, issue started: 1:30, slank notification: 1:35, escalated to DBA team: 6:23, resolved: 8:50-8:55 I have another meeting and am leaving this one. Unfortunately, I can't offer much support here. Best regards, Géraldine

Sure, thank you Ricci-Rubera, Geraldine bisclxpbisjb01.bisnodech.local

  1. Issue overview: Primary issue: Piccard application database experienced disk space exhaustion, leading to database locking sessions and severe performance degradation. The database ran out of disk space, preventing normal operations and causing application slowdowns and blockages. Although additional disk space was subsequently added, residual locking sessions in the database continued to impact performance. These locking sessions were identified as Hibernate sessions not closing correctly at the application layer, requiring application-level intervention.
  2. Timeline (as discussed / confirmed) ~00:45 CET First disk space warning received (20% remaining). ~01:3001:35 CET Disk became full; Splunk alert at ~01:35 CET confirmed 100% disk usage. ~06:23 CET Issue escalated to DBA team. 08:50 CET Additional disk space (~50 GB) added to the database (confirmed multiple times by Paweł/Adrian, also reflected in meeting chat). Post disk extension: Database became accessible again, but locking sessions persisted, causing degraded performance. (Exact issue start time for business impact aligned tentatively to 21 Apr 2026 ~19:35 EST, per meeting chat.)
  3. Customer / Business Impact Switzerland business users heavily impacted; Piccard is widely used by most Swiss business units. Key impacts reported: Users could log in, but: Navigation extremely slow. Every click took significant time. Saving changes mostly hung indefinitely (loading cursor, no completion). Debt reports and Trade Register notifications could not be loaded earlier in the day. Import / shop jobs were not functioning during the incident window. User confirmation (Daniela Pispico): Application technically accessible. Performance “very, very slow”. Not usable for productive work. Business users confirmed they were still blocked operationally despite disk extension.
  4. Current Technical State (at time of discussion) Disk space issue mitigated temporarily by adding capacity. Database still affected by: Blocking sessions. Longrunning Hibernate sessions not closing. Application performance remains degraded. ⚠️ Working copy replication not running and intentionally deferred to avoid further business impact during working hours. Risk identified: Disk space growth has been persistently high (>80%) for some time, not a oneoff condition. Sudden growth may also be linked to logs, dumps, or background processes, but no final RCA yet.
  5. Key Technical Discussions / Decisions Restart of Tomcat process identified as the next critical remediation step: Purpose: Clear blocking sessions and stuck Hibernate connections. Clarified explicitly that this is a Tomcat process restart, not a full server reboot. Confirmed: Restart should not negatively impact other applications. Not all Piccard-related services are isolated, but risk is acceptable and understood. Linux / AO Engineering involvement required to: Prepare change. Execute controlled Tomcat restart. Database replication: Acknowledged as necessary. Agreed to postpone replication until end of business day to avoid further load and risk during peak usage. Disk space must be closely monitored before replication is resumed.
  6. Actions Being Taken (with ownership) Immediate / In Progress Add disk space to stabilize database Completed at ~08:50 CET. Owner: Paweł (DBA) Prepare and execute Tomcat process restart To clear blocking and Hibernate sessions. Owners: Martin Woerner identify affected Tomcat instance(s). Oleksandr Lozinskyi / AO Linux team prepare change and restart. Status: Pending / being prepared during the meeting. Validate application behaviour postrestart Confirm: Navigation speed. Ability to save changes. Job execution. Owners: Business users (Daniela / Alessandro / DNO team). IT Dev Switzerland. NearTerm / Planned Resume database replication Planned post business hours (around end of day). Condition: Ensure sufficient disk space and stability. Owners: DBA + IT Dev Switzerland Monitor disk usage closely During the rest of the day and overnight while replication/jobs run. Owners: DBA / Engineering Followup / Preventive Disk space monitoring and alerting review Discussion raised that: Diskfull incidents are preventable and unacceptable operationally. Proposal to raise incidents or bridges when disk reaches 9899% usage. To be taken offline as a postincident / peer review action. Owners: Engineering / DBA leadership (to be defined)
  7. Open Points / Risks Root cause of: Rapid disk consumption. Longstanding high utilization trend. Need confirmation whether: Logs, dumps, backups, or specific jobs caused the sudden growth. Risk of disk filling up again once replication and jobs restart if not addressed structurally.
  8. Overall Status (at meeting point) Incident not yet resolved. Service is in a degraded, partially usable state. Restoration path agreed: Clear blocking via Tomcat restart. Validate performance with business. Resume replication safely. ECC bridge to remain active until performance and usability are confirmed restored.

CHG0136776

Ill add this incident to the master problem we have open already [olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl status tomcat.service Wed Apr 22 10:04:21 CEST 2026 ● tomcat.service - Apache Tomcat Web Application Loaded: loaded (/etc/systemd/system/tomcat.service; enabled; vendor preset: disabled) Active: active (running) since Tue 2026-04-21 15:13:25 CEST; 18h ago Process: 11443 ExecStop=/usr/share/tomcat/bin/shutdown.sh (code=exited, status=0/SUCCESS) Process: 11526 ExecStart=/usr/share/tomcat/bin/startup.sh (code=exited, status=0/SUCCESS) Main PID: 11539 (java) CGroup: /system.slice/tomcat.service └─11539 /usr/bin/java -Djava.util.logging.config.file=/usr/share/tomcat/conf/logging.properties -Djava.util.logging.manager=org.apache.... Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Starting Apache Tomcat Web Application... Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application. [olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl restart tomcat.service Wed Apr 22 10:04:55 CEST 2026 [olelozloc@bisclxpbisjb01 ~]$ date && sudo systemctl status tomcat.service Wed Apr 22 10:05:04 CEST 2026 ● tomcat.service - Apache Tomcat Web Application Loaded: loaded (/etc/systemd/system/tomcat.service; enabled; vendor preset: disabled) Active: active (running) since Wed 2026-04-22 10:05:01 CEST; 3s ago Process: 18088 ExecStop=/usr/share/tomcat/bin/shutdown.sh (code=exited, status=0/SUCCESS) Process: 18201 ExecStart=/usr/share/tomcat/bin/startup.sh (code=exited, status=0/SUCCESS) Main PID: 18218 (java) CGroup: /system.slice/tomcat.service └─18218 /usr/bin/java -Djava.util.logging.config.file=/usr/share/tomcat/conf/logging.properties -Djava.util.logging.manager=org.apache.juli.Class... Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local systemd[1]: Starting Apache Tomcat Web Application... Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local startup.sh[18201]: Existing PID file found during start. Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local startup.sh[18201]: Removing/clearing stale PID file. Apr 22 10:05:01 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application. [olelozloc@bisclxpbisjb01 ~]$

Apr 21 15:13:25 bisclxpbisjb01.bisnodech.local systemd[1]: Started Apache Tomcat Web Application. [olelozloc@bisclxpbisjb01 ~]$ Tomcat restart completed no locks on database side Tentative restoration time 4/22/2026 10:05 AM CET / 4:05 AM EST Audit database: normal state - no locks, 25 connection from 'Collecting_Services' Audit database: RED state: 6-7 locks, ~70 connection from 'Collecting_Services' Yesterday we had restart of tomcat service on same server but related to another incident INC0798444, used change CHG0136652

INC0798320 -> Piccard Score is not working; we cannot perform the outsourcing (the same as yesterday). Please fix this as soon as possible and let us know. Thank you. related issue INC0798444

INC0797772 INC0797647 INC0798267 INC0797415 INC0797747 INC0797757 INC0797445 INC0797682 INC0797490- > (Splunk Observability) Disk = 100% [H:] chsqlbkstoprw01.bisnodech.local - system backupowy się zapchał i przenoszę dane do S3 na OVH

Wczorajsze INC0798360 INC0798393-> (Splunk Observability) Disk > 90% [I:] CHPICBUDB0NPW02 przeniosłem dane

INC0795723 INC0797683 INC0797760 INC0797672 INC0797760 INC0797696 INC0797768 INC0796110 -> (Splunk Observability) Disk > 90% [M:] CHPICBUDB0NPW01 - nie było konieczności powiększyłem dysk o 50 ponieważ nie mogłem odtworzyć bazy piccard

INC0798188 INC0798497 INC0797671 INC0798216-> (Splunk Observability) Disk > 90% [P:] CHPICBUDB0PRW01 - wykonałem shrink logu dla bazy Sync na piccard

INC0798748 INC0798753- (Splunk Observability) Disk > 80% [J:] CHPICBUDB0PRW01 - nie potrzeba nic robić bo to tempdb

INC0797669 -> (Splunk Observability) Disk = 100% [P:] CHPICBUDB0PRW01 - shrink databaze

INC0798753 - > (Splunk Observability) Disk = 100% [J:] CHPICBUDB0PRW01- disk resized but waiting for enable replication

SCTASK0724812 -> Enable Sync and feed job

Today tasks SCTASK0722293 -> Fulfillment task for Request Database Support - nie mam dostępu do bazy danych i musiałem odtworzyć z ostatniego backupu jaki mam, ale

16370274 16370274 11516947 10124205 7985051 6556451 4771397 3613868 2977527 2387681 1543699 862179