Conda Environment and Package Management
Delete an existing environment (ensure it is deactivated first):
conda remove -n <env_name> --allList all available environments:
conda env list
# or
conda info --envsCreate a new environment with a specific Python version:
conda create -n <env_name> python=<version>Activate and deactivate environments:
conda activate <env_name>
conda deactivateView installed packages in the current environment:
pip listInstall packages individually or from a requirements file:
pip install <package_name>
pip install -r <path/to/requirements.txt>Integrating Conda with PyCharm
To use a newly created Conda environment in PyCharm, create a new project and configure the interpreter. Navigate to the Conda installation directory and select the conda.exe executable. PyCharm will then list the available Conda environments for selection.
To remove an interpreter from the list, go to File -> Settings -> Project:... -> Python Interpreter, click the gear icon, select Show All..., and choose Remove Interpreter.
Adjusting Pandas Display Settings
To prevent Pandas DataFrames from truncating rows or columns during output:
import pandas as pd
# Display all columns and rows
pd.set_option('display.max_columns', None)
pd.set_option('display.max_rows', None)
# Increase column width (default is 50)
pd.set_option('max_colwidth', 200)
# Prevent automatic line wrapping
pd.set_option('expand_frame_repr', False)The Mutable Default Argument Trap
Assigning a mutable object (like a list or dictionary) as a default argument value in Python leads to unexpected behavior because the default value is evaluated only once at function definition time, persisting across subsequent calls.
def append_item(val, storage=[]):
storage.append(val)
return storage
print(append_item(10)) # Output: [10]
print(append_item(20)) # Output: [10, 20]
print(append_item(30)) # Output: [10, 20, 30]The correct approach is to use None as the default value and initialize the mutable object inside the function body:
def append_item(val, storage=None):
if storage is None:
storage = []
storage.append(val)
return storageServer-Side Conda Environments and Startup Config
To create a Conda environment in a custom directory on a server:
conda create --prefix=/home/<username>/.conda/envs/<env_name> python=<version>If unable to delete an old environment because it is active, check the ~/.bashrc file. The Conda initialization block may contain a hard-coded conda activate <old_env_name> command. Modify it to the new environment or remove it, then deactivate and delete the old environment.
Essential Linux Directory Commands
Check disk usage for users:
# All users
sudo du -sh /home/*
# Current user
sudo du -sh /home/<username>Common directory operations:
ls -a: List files, including hidden ones.cd <path>: Change directory.pwd -P: Print the current working directory path.mkdir -p <path/name>: Create directories recursively.rm -rf <path/name>: Force delete files or directories recursively.
Restricting Home Directory Access
To prevent other users from accessing your home directory (does not apply to root):
sudo chmod 0700 /home/<username>Handling Division by Zero in NumPy
To safely divide arrays and replace division-by-zero results with zero:
import numpy as np
result = np.divide(numerator, denominator, out=np.zeros_like(denominator), where=denominator!=0)Detecting Anomalies in NumPy Arrays
Check for invalid values (NaN or Inf):
# Check if all values are finite
is_valid = np.isfinite(arr_data).all()
# Check for existence of NaN or Inf
has_nan = np.isnan(arr_data).any()
has_inf = np.isinf(arr_data).any()
# Locate indices of NaN values
nan_indices = np.where(np.isnan(arr_data))SciPy Skewness Output Quirk
When using scipy.stats.skew, the function returns NaN if all values in the input array are identical, which can propagate missing values into subsequent calculations.
Proper Array Indexing for Rows and Columns
Extracting specific rows and columns requires chaining index operations, not passing them simultaneously in a single bracket.
row_ids = [0, 2]
col_ids = [1, 3]
# Incorrect: extracts elements at (0,1) and (2,3)
subset_wrong = matrix[row_ids, col_ids]
# Correct: extracts the sub-matrix
subset_right = matrix[row_ids][:, col_ids]If a ValueError: ndarray is not C-contiguous occurs during further processing, enforce C-contiguous memory layout:
contiguous_arr = np.ascontiguousarray(subset_right)
# or
contiguous_arr = np.array(subset_right, order='C')Graceful Process Termination
To exit a Python script immediately without throwing an exception:
import sys
sys.exit(0)Slicing Odd or Even Columns
Use step slicing to extract every other column:
# Extract even-indexed columns
even_cols = arr[:, ::2]
# Extract odd-indexed columns
odd_cols = arr[:, 1::2]Suppressing Logarithm of Zero Warnings
When calculating entropy or similar metrics, multiplying by zero and taking the log of zero triggers RuntimeWarning. While np.where can bypass calculations, adding a tiny epsilon is often simpler for vectorized operations:
probabilities += 1e-9 # Prevent log(0)
entropy = -np.sum(probabilities * np.log(probabilities), axis=1)Extracting DataFrame Data to NumPy
To select specific columns and rows from a DataFrame based on lists of names and indices, converting directly to an array:
extracted_array = df[target_columns].iloc[row_indices].to_numpy()PyTorch Geometric Data Tensor Requirement
When constructing a PyG Data object for use with a DataLoader, node features x must be a PyTorch tensor. If passed as a NumPy array or standard list, the DataLoader misinterprets the batch dimension, often reflecting the batch size instead of the feature dimension, without throwing an explicit error at creation. Always ensure conversion:
import torch
from torch_geometric.data import Data
data_obj = Data(x=torch.tensor(node_features, dtype=torch.float), edge_index=edge_index, y=labels)