# Weekly On-call Summary Oct 11th - Oct 18th, 2021

**URL:** <https://talk.harmony.one/t/weekly-on-call-summary-oct-11th-oct-18th-2021/4578>\
**Category:** Team Ops\
**Created:** [October 19, 2021, 4:46pm UTC](https://talk.harmony.one/t/weekly-on-call-summary-oct-11th-oct-18th-2021/4578 "2021-10-19T16:46:12Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![rongjian](https://yyz1.discourse-cdn.com/flex035/user_avatar/talk.harmony.one/rongjian/32/43_2.png) [@rongjian](https://talk.harmony.one/u/rongjian)\
**Post date:** [October 19, 2021, 4:46pm UTC](https://talk.harmony.one/t/weekly-on-call-summary-oct-11th-oct-18th-2021/4578/1 "2021-10-19T16:46:13Z")

</div>

# Weekly OnCall Rotation

Oncall: RJ, Jenya (rongjian@harmony.one, jenya@harmony.one)

Duration: 10/11 08:30am - 10/18 08:30pm

# Summary

- MultiSig service is not stable and keeps going down. Mostly a simple restart bring back the service up. Root cause is being investigated and fixed by Jenya and Jack. Mostly related to aws instances and more profiling is needed.

- Migrated Multisig Postgres database from local docker to AWS cloud. It fixed the CPU issue and now the API is stable.

- Multisig indexers behind for 24 hours of data. Tracing RPC is slow and we only index 20 blocks (40 seconds of data) in ~25 seconds now. When I try to fire more requests, RPC starts to return an empty response or disconnect by timeout. Hence currently the latest actions are not visible on the UI.

- API and indexers located in different docker containers, so it will be easy to separate them. Not urgent as indexers consume very small amount of CPU/RAM

- Soph updated all validator node S1/S2/S3 disks to 500GB, now all are at 50% disk space usage. For S0, some of them are reaching 80% but nothing much we can do there since those are attached disk (not EBS)

- Auto-resolved:

# Details

### 10/12

- MultiSig Backend ([https://multisig.t.hmny.io/api/v1/about/](https://multisig.t.hmny.io/api/v1/about/) ). It was down for 59 minutes and 19 seconds. [https://harmonyone.pagerduty.com/incidents/PGVAP6P](https://harmonyone.pagerduty.com/incidents/PGVAP6P) - out of resource: need access to grant more resource
- Disk out of space: [https://harmonyone.pagerduty.com/incidents/PR01PNL](https://harmonyone.pagerduty.com/incidents/PR01PNL)
- Beacon out of sync (auto-resolved): [https://harmonyone.pagerduty.com/incidents/P9UCRE7](https://harmonyone.pagerduty.com/incidents/P9UCRE7)

10/13

- [the cpu usage rate of the mainnet service node is abnormal](https://harmonyone.pagerduty.com/incidents/Q2JDULADQI0BLE)
- beacon out of sync

10/14

- Average UnHealthyHostCount GreaterThanOrEqualToThreshold 1.0 [https://harmonyone.pagerduty.com/incidents/Q0YCQ6IOHJV597](https://harmonyone.pagerduty.com/incidents/Q0YCQ6IOHJV597)

- Monitor is DOWN: explorer-api1 ( [http://api1.explorer.t.hmny.io:3000/v0/shard/0/block/number/0](http://api1.explorer.t.hmny.io:3000/v0/shard/0/block/number/0) ) [https://harmonyone.pagerduty.com/incidents/Q2O8ID0P4M2U4U](https://harmonyone.pagerduty.com/incidents/Q2O8ID0P4M2U4U)

- Monitor is DOWN: explorer-api1-WebSocket [https://harmonyone.pagerduty.com/incidents/Q3EKRU9PZ7JPHG](https://harmonyone.pagerduty.com/incidents/Q3EKRU9PZ7JPHG)

- 
# the cpu usage rate of the mainnet service node is abnormal [https://harmonyone.pagerduty.com/incidents/Q366Q8IWGJWT6R](https://harmonyone.pagerduty.com/incidents/Q366Q8IWGJWT6R)

10/15

- MultiSig Backend ( [https://multisig.t.hmny.io/api/v1/about/](https://multisig.t.hmny.io/api/v1/about/)) [https://harmonyone.pagerduty.com/incidents/Q21DQSQXTY9TED](https://harmonyone.pagerduty.com/incidents/Q21DQSQXTY9TED)
  - Quote Jack “So, we managed to get it back up and running, but few things left to do - the healthcheck is not able to sense that the service is actually up (still says failed, don’t kill it please!), probably due to missing docker IP to instance IP mapping - the docker container startup script needs to be registered as a new service bigger picture-wise, we need to… - either commit more resources to take care of our services like these - or outsource them to like a Gnosis partner”

- Node Memory Usage Rate Alert - the memory usage rate of the mainnet shard3 node is abnormal [https://harmonyone.pagerduty.com/incidents/Q13YL3GLDLK4XM](https://harmonyone.pagerduty.com/incidents/Q13YL3GLDLK4XM)

10/16

- [Insufficient sign power of Shard1! - mainnet](https://harmonyone.pagerduty.com/incidents/Q3ZRP5ZQR22KF9)
- A few beacon out of sync issues that got auto-resolved.

10/17

- Soph updated all validator node S1/S2/S3 disks to 500GB, now all are at 50% disk space usage. For S0, some of them are reaching 80% but nothing much we can do there since those are attached disk (not EBS)
- Multi-sig server down: [https://harmonyone.pagerduty.com/incidents/Q029FY5VFDXMZ2](https://harmonyone.pagerduty.com/incidents/Q029FY5VFDXMZ2)
- Node Disk Space Alert - the free space of the mainnet shard2 node abnormal
- A few beacon out of sync issues that got auto-resolved.

10/18

- Tons of [Average UnHealthyHostCount GreaterThanOrEqualToThreshold 1.0](https://harmonyone.pagerduty.com/incidents/Q3FWESKWCOG0IJ) (a few every hour) that gets auto-resolved.
- Monitor is DOWN: MultiSig Backend [https://harmonyone.pagerduty.com/incidents/Q2TK7JGUQF2UBL](https://harmonyone.pagerduty.com/incidents/Q2TK7JGUQF2UBL)

Suggestions:

1. For each alert, link the runbook to it in the alert message.
2. For each existing host, put some description on what it’s used for and whether it’s customer-impacting.
3. Having a run through with the team on the architecture/runbook of each service
4. Expose vaultpass password in lastpass and share with have instruction on how to retrieve it
