13.6 Observability and Diagnostics
Record the reproduction time and symptoms, then examine server state, sessions and statements,
storage and memory, and relevant logs on the same timeline. M$ metadata tables describe schemas;
V$ virtual tables expose current state.
Diagnostics and logs
- Record when the problem started and ended, along with client information.
- Check server responsiveness with
machadmin -e. - Find relevant operations in
V$SESSIONandV$STMT. - Inspect virtual tables for the affected area, such as storage, memory, or ROLLUP.
- Compare server, client, and loader logs from the same time.
- Save current configuration properties and baselines before making changes.
Trace log configuration
Adjust trace levels, file sizes, and retention only as needed for diagnosis. Check property names, allowed values, and restart requirements for the current release in the Configuration Reference. Logs may contain sensitive SQL and data, so set access permissions and retention periods.
Server logs
The default trace directory is $MACHBASE_HOME/trc. Inspect the actual directory and settings
instead of assuming fixed filenames.
ls -lh "$MACHBASE_HOME/trc"
tail -n 200 "$MACHBASE_HOME/trc/machbase.trc"Review the first error, preceding warnings, server startup and shutdown, and checkpoint and storage events chronologically, rather than merely counting error strings. Do not manually bulk-delete log files.
machsql logs
machsql.history may contain credentials or sensitive SQL. Restrict history file permissions
for production accounts and review the contents before sharing diagnostics. Preserve reproduction
SQL together with the target database, execution time, results, and errors.
machloader logs
For bulk loading, check the process exit code, summary, logs, and rejected-row file together.
- Schema and input column count and order
- Delimiters, quote characters, and encoding
- NULL and DATETIME formats
- First failed row and recurring error codes
- Final successful and failed row counts
Do not blindly replay the entire rejected-row file. Fix the cause, validate a sample, then reprocess only the failed rows.
Metadata tables
Inspect the current database schema through M$SYS_TABLES, M$SYS_COLUMNS, M$SYS_INDEXES,
and related tables. Include database, owner, and object IDs in joins instead of joining by
name alone. Avoid application dependencies on reserved object names or internal table structures.
Virtual tables
First check the actual columns in the current release.
SELECT * FROM V$SESSION LIMIT 1;
SELECT * FROM V$STMT LIMIT 1;
SELECT * FROM V$PROPERTY LIMIT 1;
SELECT * FROM V$STORAGE_USAGE LIMIT 1;
SELECT * FROM V$SYSMEM LIMIT 1;
SELECT * FROM V$ROLLUP LIMIT 1;
SELECT * FROM V$LICENSE_INFO LIMIT 1;For the complete list and column definitions, see the System Catalog.
Monitoring and capacity management
Base alerts on normal baselines, growth rates, peak workloads, and recovery headroom rather than fixed thresholds. Review file system and database storage usage together, including separate space used by backups and exports.
Server state
"$MACHBASE_HOME/bin/machadmin" -eA running process alone does not prove server health. Also check a native connection, a lightweight SQL query, and recent server logs.
Sessions and running SQL
SELECT id, user_name, user_ip, login_time, client_type
FROM V$SESSION
ORDER BY login_time DESC;
SELECT sess_id, id AS stmt_id, state, record_size, query
FROM V$STMT
ORDER BY sess_id, id;A long-running operation is not necessarily an error. Check the workload type, processed rows, client timeouts, and I/O and CPU state before deciding to cancel or kill it.
Disk capacity
df -h "$MACHBASE_HOME"
du -sh "$MACHBASE_HOME/dbs"If a separate DBS_PATH is configured, check the actual path. Do not directly edit or delete
internal partition files.
Memory
Compare OS available memory and swap with V$SYSMEM, cache usage, and query concurrency.
Do not make automation depend on internal manager names.
Backup validation
Do more than check backup command success. Record the path, size, and completion status, then mount or restore in an isolated environment and verify key tables, row counts, time ranges, and sample queries.
Collect diagnostic information
- Release and edition
- Incident time and time zone
- Reproduction commands, database, and user
- Server, client, and loader logs
- Relevant virtual table results
- OS CPU, I/O, memory, and disk metrics
- Recent schema, configuration, or deployment changes
- Actions already attempted and their results
Remove credentials, AUTH KEYs, personal information, and sensitive raw data from support materials.