# HIK sync reliability review and prevention plan Reviewed: 2026-09-24. Scope: current workspace code and the three supplied SQL exports. SQL exports were parsed as data, not executed. No production data or application code was changed. ## Assessment Event ingestion was operating at the time represented by the exports, but end-to-end completeness is not established. The latest event is dated `2026-09-24 14:44:25`, received at `14:44:55` (database values). A successful receiver response or health check does not establish that punches appear in the frontend. | Export check | Result | | --- | ---: | | Raw HIK events | 10,626 | | Employee mappings | 190 | | Biometric logs | 18,527 | | Events with no mirrored_at | 98 across 30 device/employee pairs | | Pending events with no mirror_error | All 98 | | Events without an active device/employee mapping in the supplied map snapshot | 4,332 | | Logs whose person_id has no active mapping in the supplied map snapshot | 6,716 across 66 person IDs | | Those logs by source | 4,235 hikvision_api; 2,481 ivms_csv | | Logs missing company_id or brand_id | 2,430 | | Maps missing company_id or brand_id | 1 (device employee number 1) | | Events marked mirrored with an active map but no matching mapped person_id/time_log in the log export | 57 | These are snapshot comparisons, not proof that every unmatched record needs a new map. CSV records do not necessarily need HIK maps. Employees/users were not supplied, so active employment, canonical identity, and correct company/brand cannot be validated. The mapping export header is one minute earlier than the other exports; concurrent changes, legacy identities, and historical cleanup can also affect comparisons. The 57 unmatched mirrored events need investigation before any recovery writes. Header timestamps and row timestamps must not be assumed to use the same timezone. ## Confirmed code findings 1. **Frontend visibility is not directly joined to the HIK mapping table.** `frontend/src/components/biometrics/BiometricsRecords.jsx` calls `fetch_records`. `backend/attendance_biometrics/fetch_records.php` requires an active matching employee, the selected employee company/brand, and matching log company/brand. Thus the 2,430 unscoped logs cannot pass the current scoped query. Populating maps alone does not update those existing logs in this code. The exact SQL previously used and employee data would be needed to establish the original incident's full cause. 2. **Automatic mapping already exists.** `backend/hik-sync/lib.php::ensureEventEmployeeMap()` can insert a missing map when the device employee number uniquely matches an active employee's biometric number, falling back to employee_id only when the biometric number is empty. It deliberately does not match by name. `add_employee.php`, `bulk_add_employee.php`, and `update_employee.php` also call `solidmark_hik_map_sync()`. 3. **An active map is treated as sufficient before validating scope.** `ensureEventEmployeeMap()` returns early for any active map. `mirrorEvent()` subsequently requires matching company/brand and active employee identity. A stale or unscoped active map can therefore block mirroring without repair. 4. **Unresolved events can be silent.** If employee resolution or the mirror join finds no row, mirroring returns false. The receiver and replay scripts record mirror_error only for exceptions. `replay.php` can report ok=true with unresolved events remaining. Receiver `unmapped_or_pending = stored - mirrored` is not an accurate pending count for duplicate/retried events. 5. **Repair coverage is incomplete.** `repair_missing_maps.php` checks whether an employee has any active map, not whether every required device has a valid scoped map. One map on another device, or a map with stale scope, can mask the gap. 6. **Device configuration resolution differs.** The receiver supports `HORIZON_HIK_CONFIG_PATH` and the ProgramData configuration. Employee map synchronization reads only the repository config plus devices already present in the database. A newly configured external device may not be provisioned until events arrive. 7. **Replay exists but automatic scheduling was not established.** No scheduler wiring for replay/repair was found in the searched backend/docs files. The host scheduler was not inspected. Saving an employee does not itself demonstrate that older pending events are retried. 8. **Health checks only database connectivity.** `health.php` does SELECT 1; it does not check mapping failures, pending age, or mirror completeness. 9. **Production parity needs verification.** The receiver README describes this directory as a reference package and warns that the live receiver may have newer resolution/replay logic. Its referenced security-merge document is absent at the stated path. Findings describe this checkout, not an independently inspected deployed receiver. ## Prevention blueprint ### 1. Database architecture and ERD Preserve raw punches even if their employee cannot yet be resolved. Match stable IDs, not names. Missing or ambiguous identities enter a review queue; do not manufacture employees or assign a default brand. | Entity | Keys / relationship | Tracking and proposed changes | | --- | --- | --- | | employees | Existing employee_id; validated biometric_employee_no | Authoritative current identity and scope; inspect existing constraints before adding uniqueness | | hik_sync_employee_map | Existing unique (device_key, hik_employee_no); logical link to employee identity | Retain created_at/updated_at; validate active status, identity and scope; preserve legacy aliases | | hik_sync_events | Existing event_key and device identity | Retain received_at, mirrored_at, mirror_error; add resolution_status, attempt_count, last_attempt_at, next_retry_at and resolved employee reference after schema review | | tbl_biometrics_logs | Existing id; logical link to resolved employee | Review and apply the existing source-identity migration: source_type/source_event_key, with the intended unique constraint verified against historical duplicates; preserve company/brand scope | | hik_sync_audit (proposed) | id PK; event_key/map_id references; actor reference or service identity | action, reason, before_json, after_json, request/job ID, company/brand, created_at; append-only | Logical ERD: `employees 1 -> N employee mappings`; `(device_key, hik_employee_no) mapping 1 -> N events`; `event -> 0..1 source-linked log`; `mapping/event -> N audit entries`. Legacy time/person deduplication may associate multiple source events with one existing punch; preserve that distinction explicitly rather than claiming each event inserted a row. Unmapped events must be allowed to exist without a mapping foreign key. Index pending retry scans by resolution_status/next_retry_at/id and device/employee lookups by device_key/hik_employee_no. Check existing indexes and query plans before adding indexes. Keep IDs textual, preserving leading zeros. Scope changes for historical punches require explicit policy; never automatically move all historical attendance to an employee's current brand. ### 2. RESTful PHP API and worker plan Proposed admin routes follow the existing `/api/attendance_biometrics/` convention; these routes are not implemented by this review. | Method / route | Purpose | Payload / response | | --- | --- | --- | | POST existing receiver `backend/hik-sync/receive.php` | Persist signed device events; attempt safe mapping/mirroring | Existing device_key/events payload; response separates stored, duplicate, mirrored, already_mirrored, unresolved and failed counts based on actual final event state | | GET `/api/attendance_biometrics/sync_status` | Authorized sync health | Per-device last_received_at, pending_count, oldest_pending_at, unresolved_count, failure_count; freshness threshold configurable by device working schedule | | GET `/api/attendance_biometrics/sync_issues` | Paginated unresolved queue | cursor, limit (max 100), device_key, reason, from/to, employee query, allowlisted sort; returns items, next_cursor, total if inexpensive | | POST `/api/attendance_biometrics/resolve_sync_issue` | Resolve verified identity/scope conflicts | event/map identifier, selected employee_id, expected version, reason; returns resolution and queued retry status; reject stale edits with 409 | | POST `/api/attendance_biometrics/retry_sync` | Queue bounded targeted retries | Device/employee identifiers or explicit event IDs; returns 202/job_id, not a synchronous unbounded replay | Standard admin errors: `{ "ok": false, "error": { "code": "IDENTITY_AMBIGUOUS", "message": "Select a verified employee.", "fields": {} }, "request_id": "..." }`. Use 400/422 for invalid data, 401 for unauthenticated access, 403 for unauthorized scope, 404 for inaccessible/missing records, 409 for conflicts, and 500 for unexpected failures. Preserve receiver compatibility when extending its existing error shape. Use session authentication and server-enforced RBAC/scope for admin APIs, CSRF protection on session mutations, prepared statements, allowlisted sorting, and private exception logging. Keep receiver HMAC/device authentication separate. Never expose raw payloads or unassigned cross-brand identities to ordinary brand users; genuinely unscoped records require an authorized central administrator. Worker behavior: 1. Share configuration resolution between receiver and employee synchronization. 2. Reconcile employee/device pairs, including existing maps with invalid scope. Auto-repair only uniquely verified identities; retain conflicts for review and do not reactivate deliberately disabled aliases without policy. 3. Save employee/mapping updates atomically and queue targeted retries after commit. 4. Run a bounded scheduled retry worker (proposed initial cadence: every five minutes) with locking, backoff and cursor batches. Recheck unresolved identities without starving newer events. Confirm deployment paths and database before registering the schedule. 5. Record explicit reasons: missing_employee, ambiguous_identity, inactive_employee, invalid_scope, mapping_conflict, database_error. A successfully accepted event is not necessarily successfully mirrored. 6. Reconcile marked-mirrored events against source-linked logs separately from ordinary pending replay. Investigate the 57 snapshot mismatches before applying recovery. Do not blindly reset mirrored_at. 7. Alert authorized operators on aged pending events or repeated failures. The existing connectivity endpoint may remain a simple liveness check; detailed readiness belongs in authenticated monitoring. ### 3. React frontend flow Extend the biometrics page with `SyncStatusSummary -> SyncIssuesTable -> ResolutionDialog`. Keep existing attendance views scoped. Fetch paginated issues independently from attendance records, debounce ID/name search, cancel stale requests, and refresh status after resolution. Use stable row IDs; avoid loading the full event history into React. Introduce virtualization only if bounded pages still require it. Use mobile-first stacked issue cards and a desktop table. Include loading, empty, partial-failure and retry states. Resolution modal displays device number, verified candidate identity, scope, reason and affected event count; require a reason and validate the selected employee on the server. Trap focus, label controls, restore focus after closing, and use a confirmation step for identity reassignment. Show unresolved counts clearly without exposing cross-brand records. Central administrators handle unscoped records; authorized brand administrators handle only verified in-scope records. ### 4. Audit and compliance strategy Append an audit entry for automatic mapping creation, scope repair, manual reassignment, retry/recovery decisions and critical failures. Capture actor/service identity, old/new state, timestamp, device/event/map keys, company/brand, reason and correlation/job ID. Write mutation audits in the same transaction as the change. Grant the runtime audit role insert/read permissions without update/delete; restrict maintenance separately. Do not copy device secrets, credentials, picture URLs or full biometric raw payloads into audit JSON. Define retention and access controls with the existing attendance policy. ## Validation and deployment sequence 1. Compare deployed receiver code/configuration and schema with this checkout. Obtain a consistent snapshot including employees and users; rerun read-only checks using both canonical and supported legacy identity relationships. 2. Confirm the correct company/brand and identity for pending/unscoped records. Inspect the original repair query if available. Do not infer identity from matching names. 3. Implement shared resolution, explicit pending reasons and accurate counters first. Then implement scoped reconciliation, targeted scheduled retries, auditing and the admin queue. 4. Use an isolated database fixture to test: unique new employee mapping, already-active unscoped map, conflicting identities, intentionally inactive maps, leading-zero IDs, external-config-only devices, one missing map across two devices, duplicate delivery, concurrent retries, transaction rollback, and successful replay after employee correction. Assert both log persistence and authenticated frontend visibility; assert cross-brand isolation. 5. Dry-run reconciliation against a production snapshot, review proposed changes, then deploy migrations/code through the normal release process and verify scheduler execution. Retain recoverable backups before data repair. This review does not certify current live health or deploy the proposed changes. Prevention requires verified employee identity, correct log scope, automatic retry, and visible exceptions together; merely repopulating the map table is insufficient. ## Horizon HIK Agent review (v2.5.6 source) Reviewed source: `C:\Users\centr\OneDrive\Desktop\Horizon-HIK-Agent`. No files in that project were changed. ### Current failure and resume behavior The agent already provides a durable per-device checkpoint and safe continuation for ordinary ISAPI failures. ```mermaid flowchart LR A[Poll ISAPI AcsEvent] -->|Query succeeds| B[Save accepted events to local pending-events outbox] B -->|Each event file is durable| C[Persist per-device serial number and event time] C --> D[Send outbox to Horizon receiver] A -->|Timeout, authentication, HTTP or invalid response| E[Log device error; retain existing checkpoint] E --> F[Next scheduled poll] F --> A D -->|Receiver unavailable| G[Retain outbox files; retry next cycle] ``` - State is stored under `%LOCALAPPDATA%\HorizonHIKAgent`, including `state.json` (per-device `lastSerialNo`, `lastEventTime`, error/success timestamps) and one JSON file per unacknowledged event in `pending-events`. - ISAPI reads use Digest authentication and `ISAPI/AccessControl/AcsEvent`. The query time starts at the last successful event time minus `queryOverlapSeconds`; it accepts only events with a serial number greater than the stored serial checkpoint. - The checkpoint advances only after every accepted event is atomically written to the local outbox. If an ISAPI read, parsing step, or outbox write fails, that device's checkpoint is unchanged. This is the required pause/continue behavior. - The agent keeps polling a failed configured device at the normal interval. Other configured devices and delivery of already queued events continue. The dashboard records each device's `lastHikError`, `lastHikErrorAt`, `lastHikSuccessAt`, and a capability cache; its logs retain the ISAPI error message and HIK HTTP-400 response detail when supplied. - Receiver delivery is also durable: events are removed from the local outbox only after an authenticated successful receiver response. The receiver separately deduplicates by device key and serial number. - If no enabled device, device credentials, receiver URL, or receiver secret are configured, `configReady()` is false. Start is rejected and automatic sync does not start after a restart. A deliberately disabled device is skipped and its stored checkpoint is preserved. ### Gaps that need correction | Priority | Gap | Effect | | --- | --- | --- | | High | The agent acknowledges and deletes a local outbox event after the receiver merely stores it, even when the receiver reports `unmapped_or_pending`. | An event missing an HRIS map no longer exists in the agent queue. Recovery then depends entirely on a server-side replay job, which is not automatically scheduled in this checkout. | | High | The receiver's `unmapped_or_pending` count is computed as `stored - mirrored`. It excludes duplicate events that remain unresolved. | The agent's displayed pending mapping count is inaccurate; repeated device delivery can look healthy while HRIS punches remain missing. | | High | The agent does not have explicit device health states or a failure backoff. | An unconfigured/invalid device or persistent ISAPI 401/404/400 is retried at full polling frequency and appears only in local logs/status. | | High | Saving configuration does not require a successful device and receiver test before automatic sync can remain enabled. A later incomplete configuration makes `syncOnce()` fail before it records a useful state/message; the scheduler quietly retries. | “Not configured” can be silent after settings are edited. | | Medium | Checkpoint correctness relies on monotonically increasing HIK serial numbers. | If a device resets/replaces its event database and serial numbers restart below the saved cursor, every new event can be rejected indefinitely. | | Medium | ISAPI failures are logged as raw error strings in the automatic path. | The dashboard can identify failure but cannot distinguish a configuration fault, temporary network outage, incompatible firmware, device busy response, invalid payload, or serial reset in a structured way. | | Medium | `consecutiveErrors` is incremented but does not drive alerting, backoff, pause, or recovery policy. | The value is informational rather than operational. | | Medium | A device disabled in the UI has no explicit recorded reason, disabled-at time, or reminder that its checkpoint is intentionally frozen. | Operators may confuse a deliberate pause with a device outage. | The 400 fallback for large batches and attendance-only filters is good compatibility handling. It retries once at a safe batch size or with all event types before treating the query as failed. It should remain distinct from a device-health failure. ### Recommended agent design Add an explicit durable state per device: `ready`, `paused_unconfigured`, `retrying`, `auth_failed`, `unsupported`, `needs_attention`, and `disabled`. Store `failure_code`, sanitized failure detail, `failure_count`, `next_attempt_at`, `last_success_at`, `last_query_start`, `last_query_end`, `last_committed_serial`, `last_committed_event_time`, and a device identity fingerprint (model/serial/firmware where available). The phrase “checkpoint” must mean the last event safely persisted to the local outbox, never the last successful HTTP call. Classify failures as follows: | Condition | Agent behavior | Checkpoint behavior | | --- | --- | --- | | Offline, timeout, DNS, connection refused, TLS transient failure, 5xx | Retry with exponential backoff and jitter; keep delivering existing outbox items | Preserve | | HTTP 401/403 or Digest failure | Pause only that device after one verified failure; require an administrator configuration/test action to resume | Preserve | | Unsupported endpoint/firmware (404/405) | Pause that device and display a remediation message | Preserve | | ISAPI HTTP 400 after safe batch/filter fallbacks, malformed XML/JSON, device busy response | Mark retrying or needs_attention depending on structured HIK substatus; capture sanitized status detail | Preserve | | Device not configured or receiver configuration incomplete | Set `paused_unconfigured`; do not schedule polling for that profile; show configuration required | Preserve | | Local outbox full or unreadable | Pause capture and raise a critical visible alert; continue no cursor advancement | Preserve | | Receiver accepts raw event but cannot map/mirror it | Delete the agent outbox only because the server durable raw-event record is confirmed; server must retain explicit unresolved status and automatically replay after a mapping change | Preserve server-side unresolved event until mirrored | Use capped exponential backoff, for example 15 seconds, 30 seconds, 60 seconds, 2 minutes, then a maximum of 10 minutes, with random jitter. Manual **Test Biometric Device** and a successful configuration save should bypass the delay. One failed device must never pause another device or delivery of events already held locally. For a suspected serial rollback, do not automatically reset the cursor. Compare device identity and event time: if the device identity changed or the returned newest event time is materially later than the saved event time while every serial is lower, mark `needs_attention` and offer an administrator-only “start a new epoch” action. That action must retain the prior checkpoint/audit trail and use the overlap window plus receiver event-key deduplication. This avoids both a permanent data gap and a blind replay of old attendance. ### Database and API additions Use the existing `hik_sync_events` table as the server's durable source of truth after receipt. Add an operational table rather than overloading employee mappings: | Table | Required columns | | --- | --- | | `hik_sync_device_health` | `device_key` PK, `state`, `failure_code`, `failure_detail`, `failure_count`, `next_attempt_at`, `last_agent_seen_at`, `last_hik_success_at`, `last_hik_error_at`, `last_committed_serial`, `last_committed_event_time`, `device_identity_json`, `updated_at` | | `hik_sync_audit` | Existing proposed immutable audit table: actor/service, action, old/new state, reason, device/event identifiers, timestamp and correlation ID | Extend `hik_sync_events` with an explicit resolution outcome (`pending_mapping`, `pending_scope`, `mirrored`, `mirror_failed`, `ambiguous_identity`) and retry metadata. Index unresolved/retry scans. The server must update this outcome for duplicate delivery as well as fresh events. Proposed routes: | Method / route | Purpose | | --- | --- | | `POST /api/attendance_biometrics/agent_heartbeat` | Authenticated agent reports device health, checkpoint and configuration state; no secrets or raw punch payloads | | `GET /api/attendance_biometrics/sync_status` | Authorized HRIS operators see device state, last success, error, checkpoint age, local/server pending counts and oldest unresolved event | | `POST /api/attendance_biometrics/retry_sync` | Authorized targeted server replay after a verified mapping or scope correction | | `POST /api/attendance_biometrics/reset_device_epoch` | Central administrator only; records approved serial-reset recovery with expected device identity and reason | Return structured receiver acknowledgements: `stored`, `duplicates`, `mirrored`, `unresolved`, `failed`, and event keys/counts for each state. Do not use `stored - mirrored` as a proxy. The agent can still remove its outbox only after the raw event has been stored idempotently, while the HRIS server owns the unresolved-mapping replay lifecycle. ### Frontend and audit flow The agent dashboard should show a compact card for every configured device: state, model/firmware, last successful ISAPI query, last checkpoint serial/time, next retry, error category and a clear action. Use a configuration-required dialog for unconfigured profiles, a read-only warning for disabled profiles, and a confirmation dialog for serial-epoch recovery. The HRIS Biometrics page should show server-side unresolved mapping counts separately from local agent delivery pending, because they mean different things. Audit configuration changes, device enable/disable actions, authentication failures, automatic pause/resume, checkpoint epoch resets, map resolutions and server replays. Never include HIK passwords, HMAC secrets, raw authorization headers, or full biometric payloads in audit/health responses. ### Validation before release Test the agent in a disposable local data directory with mocked ISAPI/receiver responses. Cover: no device configured; disabled device; bad device password; endpoint 404/405; timeout; HTTP 400 then safe fallback; malformed event response; local outbox full; receiver timeout; receiver raw-store success with unresolved mapping; mapping correction followed by server replay; engine restart after each checkpoint boundary; multiple devices where one is down; and a serial-reset scenario. Assert that every accepted event is either in the local outbox or durably present in `hik_sync_events`, and that no checkpoint advances before that condition is true. ### Deferred implementation: historical device-record recovery User decision on 2026-09-24: document this work only; do not change the Horizon HIK Agent or Horizon receiver yet. Observed business case: records remain available in the biometric device but never reach Horizon. The agent must support a controlled recovery scan for those records without changing normal live-sync behavior or creating duplicate biometric logs. Possible causes in the reviewed agent: 1. A saved per-device `lastSerialNo` excludes older device records from every later normal poll, because live polling accepts only serials greater than that checkpoint. 2. The Agent Punch Guard intentionally discards a record as a repeated punch or an already-used attendance slot. The device retains the record, but it is never queued for Horizon. 3. A device read can recover after ISAPI interruption while the agent's cursor already represents a later successfully processed event. Normal live polling then does not revisit an earlier gap. 4. The receiver can durably store a raw event but leave it unresolved because its employee mapping or company/brand scope was unavailable. Recovery requires server-side replay after the mapping is corrected. Required later implementation: | Capability | Required behavior | | --- | --- | | Automatic ISAPI recovery scan | After a device changes from an ISAPI failure state to a healthy state, query a configurable historical window before the current checkpoint. Run this as a recovery scan, separate from normal live polling. | | Manual historical re-read | Provide an administrator-only action to select one device and a start/end date/time. Show the proposed range and confirmation before querying. Do not reset the normal checkpoint merely to perform this scan. | | Idempotent recovery | Submit re-read events through the existing stable `(device_key, serial_no)` / event-key deduplication. Existing receiver events and source-linked biometric logs must not be duplicated. | | Skip-reason audit | For every device record excluded from delivery, log a structured reason: `below_checkpoint`, `punch_guard_duplicate`, `punch_guard_slot_used`, `isapi_failure`, `unconfigured`, `local_outbox_full`, `receiver_delivery_failed`, or `unresolved_employee_mapping`. | | Recovery outcome | Show queried, accepted, already-known, punch-guard-excluded, queued, delivered, unresolved, and failed counts per device and scan. Persist an audit record with actor, date range, device identity, checkpoint before/after, and reason. | | Safe cursor semantics | The normal checkpoint remains the last event safely written to the local outbox. A recovery scan must not move it backward. A serial-reset recovery must be a separate, audited administrator action. | | Unconfigured/failed device behavior | Mark the device `paused_unconfigured`, `auth_failed`, `unsupported`, or `retrying` as appropriate. Keep its checkpoint unchanged. Resume ordinary polling only after a successful device test or corrected configuration. | Recommended recovery sequence: ```mermaid flowchart TD A[ISAPI device recovers or administrator starts recovery] --> B[Choose a bounded historical time range] B --> C[Query HIK ISAPI without normal serial checkpoint exclusion] C --> D{Known event key or source log?} D -->|Yes| E[Count as already known; do not duplicate] D -->|No| F[Apply recovery policy and record any punch-guard exclusion] F --> G[Durably write accepted event to local outbox] G --> H[Send to receiver] H --> I{Mapped and scoped?} I -->|Yes| J[Mirror to biometric log] I -->|No| K[Keep server event unresolved and retry after mapping repair] ``` Before implementation, collect the affected machine's `%LOCALAPPDATA%\HorizonHIKAgent\state.json` and relevant JSONL logs. Compare its per-device checkpoint and punch-guard decisions against a bounded ISAPI export from the same period. This identifies whether the gap came from checkpoint exclusion, punch guard, receiver delivery, or employee mapping before any recovery writes are attempted.