cagraph

_images/cagraph-logo.png

Python library for generating graphs from calcium imaging data of neural activity.

Installation

Install with pip

pip install cagraph

preprocessing

cagraph

class cagraph.CaGraph(data, node_labels=None, node_metadata=None, dataset_id=None, threshold=None)

Author: Veronica Porubsky Author Github: https://github.com/vporubsky Author ORCID: https://orcid.org/0000-0001-7216-3368

Class: CaGraph(data_file, node_labels=None, node_metadata=None, dataset_id=None, threshold=None)

This class provides functionality to easily visualize time-series data of neuronal activity and to compute correlation metrics of neuronal networks, and generate graph objects which can be analyzed using graph theory. There are several graph theoretical metrics for further analysis of neuronal network connectivity patterns.

Attributes

datastr or numpy.ndarray

A string pointing to the file to be used for data analysis, or a numpy.ndarray containing data loaded into memory. The first (idx 0) row must contain timepoints, the subsequent rows each represent a single neuron timeseries of calcium fluorescence data sampled at the timepoints specified in the first row.

node_labels: list

A list of identifiers for each row of calcium imaging data (each neuron) in the data_file passed to CaGraph.

node_metadata: dict

Contains metadata which is associated with neurons in the network. Each key in the dictionary will be added as an attribute to the CaGraph object, and the associated value will be . Each value

dataset_id: str

A unique identifier can be added to the CaGraph object.

threshold: float

Sets a threshold to be used for thresholded graph.

class GraphTheory(neuron_dynamics, time, pearsons_correlation_matrix, graph, num_neurons, labels)
compare_graphs(graph1, graph2)
Returns:

get_betweenness_centrality(graph=None, return_type='list')

Returns the betweenness centrality scores for all nodes.

Parameters:
  • graph –

  • return_type –

Returns:

get_clustering_coefficient(graph=None, return_type='list')

Returns a list of clustering coefficient values for each node.

Parameters:
  • graph – networkx.Graph object

  • return_type – str

get_communities(graph=None, return_type='list')

Returns a list of communities, composed of a group of nodes.

Parameters:
  • return_type – list

  • graph – networkx.Graph object

Returns:

node_groups: list

get_connected_components(graph=None) → list

Returns connected components with more than one node.

Parameters:

graph – networkx.Graph object

Returns:

list

get_correlated_pair_ratio(graph=None, return_type='list')

Computes the number of connections each neuron has, divided by the nuber of cells in the field of view. This method is described in Jimenez et al. 2020: https://www.nature.com/articles/s41467-020-17270-w#Sec8

Parameters:
  • graph – networkx.Graph object

  • return_type – str

get_degree(graph=None, return_type='list')

Returns iterator object of (node, degree) pairs.

Parameters:
  • graph – networkx.Graph object

  • return_type – str

get_density(graph=None)

Returns the ratio of edges present in the graph out of the total possible edges.

Parameters:

graph – networkx.Graph object

Returns:

float

get_eigenvector_centrality(graph=None, return_type='list')

Compute the eigenvector centrality of all graph nodes, the measure of influence each node has on the graph.

Parameters:
  • graph – networkx.Graph object

  • return_type – str

get_hubs(graph=None, return_type='list')

Computes hub nodes using the normalized betweenness centrality scores. This method sets an outlier threshold using the inter-quartile range of the betweenness centrality score distribution and returns nodes with with scores above the outlier threshold.

Parameters:
  • graph –

  • return_type –

Returns:

get_largest_connected_component(graph=None) → Graph

Returns a subgraph containing the largest connected component.

Parameters:

graph – networkx.Graph object

Returns:

networkx.Graph object

get_path_length()

Returns the characteristic path length.

Returns:

get_smallworld_largest_subnetwork(graph=None) → float
Parameters:

graph – networkx.Graph object

Returns:

float

property graph
property neuron_dynamics
property node_labels
property num_neurons
property pearsons_correlation_matrix
property time
class Plotting(neuron_dynamics, time, pearsons_correlation_matrix, graph, num_neurons)
get_single_neuron_timecourse(neuron_trace_number) → ndarray

Return time vector stacked on the recorded calcium fluorescence for the neuron of interest.

Parameters:

neuron_trace_number – int

Returns:

numpy.ndarray

property graph
property neuron_dynamics
property num_neurons
property pearsons_correlation_matrix
plot_all_neurons_timecourse(title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
Parameters:
  • save_format –

  • title –

  • y_label –

  • x_label –

  • show_plot –

  • save_plot –

  • save_path –

  • dpi –

plot_correlation_heatmap(correlation_matrix=None, title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')

Plots a heatmap of the correlation matrix.

Parameters:
  • save_format –

  • correlation_matrix –

  • title –

  • y_label –

  • x_label –

  • show_plot –

  • save_plot –

  • save_path –

  • dpi –

Returns:

plot_multi_neuron_timecourse(neuron_trace_labels, palette=None, title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')

Plots multiple individual calcium fluorescence traces, stacked vertically.

Parameters:
  • save_format –

  • neuron_trace_labels – list

  • palette – list

  • title –

  • y_label –

  • x_label –

  • show_plot –

  • save_plot –

  • save_path –

  • dpi –

Returns:

plot_single_neuron_timecourse(neuron_trace_number, title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
Parameters:
  • save_format –

  • neuron_trace_number – int

  • title –

  • y_label –

  • x_label –

  • show_plot –

  • save_plot –

  • save_path –

  • dpi –

Returns:

property time
property betweenness_centrality
property clustering_coefficient
property communities
property correlated_pair_ratio
property data
property dataset_id
property degree
draw_graph(graph=None, position=None, node_size=25, node_color='b', alpha=0.5)

Draws a simple graph.

Parameters:
  • graph – networkx.Graph object

  • position – dict

  • node_size – int

  • node_color – str

  • alpha – float

Returns:

property dt
get_adjacency_matrix(threshold=None) → ndarray

Returns the adjacency matrix of a graph where edges exist when greater than the provided threshold.

Uses the Pearson’s correlation matrix.

Returns:

numpy.ndarray

get_erdos_renyi_graph(graph=None) → Graph

Generates an Erdos-Renyi random graph using a graph edge density metric computed from the graph to be randomized.

Parameters:

graph –

Returns:

networkx.Graph object

get_graph(threshold=None, weighted=False) → Graph

Automatically generate graph object from numpy adjacency matrix.

Parameters:
  • threshold –

  • weighted – bool

Returns:

networkx.Graph object

get_laplacian_matrix(graph=None) → ndarray

Returns the Laplacian matrix of the specified graph.

Parameters:

graph – networkx.Graph object

Returns:

get_pearsons_correlation_matrix(data_matrix=None) → ndarray

Returns the Pearson’s correlation for all neuron pairs.

A loaded numpy.ndarray dataset can be passed to the method for analysis, otherwise the dataset passed to the CaGraph object constructor will be used.

Parameters:

data_matrix – numpy.ndarray

Returns:

get_random_graph(graph=None) → Graph

Generates a random graph. The nx.algorithms.smallworld.random_reference is adapted from the Maslov and Sneppen (2002) algorithm. It randomizes the existing graph.

Returns:

networkx.Graph object

get_report(parsing_nodes=None, parse_by_attribute=None, parsing_operation=None, parsing_value=None, save_report=False, save_path=None, save_filename=None, save_filetype=None)
Parameters:
  • save_filetype –

  • save_filename –

  • save_path –

  • save_report –

  • parsing_nodes –

  • parse_by_attribute – str

  • parsing_operation – str

  • parsing_value – float

Returns:

dict

get_weight_matrix() → ndarray

Returns a weighted connectivity matrix with zero along the diagonal. No threshold is applied.

Returns:

numpy.ndarray

property graph
property hubs
static load(file_path)
Parameters:

file_path –

Returns:

property neuron_dynamics
property node_labels
property num_neurons
property pearsons_correlation_matrix
reset()

Resets the CaGraph object graph attribute to the original state at the time the object was created.

save(file_path=None)
Parameters:

file_path –

Returns:

sensitivity_analysis(data, threshold=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')

Generates a series of graphs around the recommended or user-specified threshold and shows the number of edits required to transform the original graph to the series of graphs.

If many edits are required, the graphs are dissimilar.

Parameters:
  • save_format –

  • dpi –

  • save_path –

  • save_plot –

  • data –

  • threshold –

  • show_plot –

Returns:

property threshold
property time
class cagraph.CaGraphBatch(data_path, group_id=None, threshold=None, threshold_averaged=False)

Class for running batched analyses.

Only directories can be passed to CaGraphBatch. Node metadata cannot be added to CaGraph objects in the batched analysis. Future versions will include the node_metadata attribute.

property dataset_identifiers
get_cagraph(condition_label) → CaGraph

Return a CaGraph object for the specified dataset condition_label

Parameters:

condition_label – str

Returns:

CaGraph

get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)

Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.

Parameters:
  • save_report –

  • save_path –

  • save_filename –

  • save_filetype –

Returns:

property group_id
static load(file_path)
Parameters:

file_path –

Returns:

save(file_path=None)
Parameters:

file_path –

Returns:

save_individual_dataset_reports(save_path=None, save_filetype=None)

Saves individual reports for each of the specified datasets. Individual filenames will be generated using the filename name of the dataset from which the analysis is derived.

This will result in the same analysis that can be done by creating a CaGraph object using a single dataset.

Parameters:
  • save_path – str

  • save_filetype – str (‘csv’, ‘HDF5’, ‘xlsx’)

Returns:

property threshold
class cagraph.CaGraphBatchTimeSamples(data_path, group_id=None, time_samples=None, condition_labels=None, threshold=None, threshold_averaged=False)

Class for running batched analyses on datasets that have distinct time periods to separate into samples.

Only directories can be passed to CaGraphBatchTimeSamples. Node metadata cannot be added to CaGraph objects in the batched analysis. Future versions will include the node_metadata attribute.

property dataset_identifiers
get_cagraph(condition_label) → CaGraph

Return a CaGraph object for the specified dataset condition_label

Parameters:

condition_label – str

Returns:

CaGraph

get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)

Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.

Parameters:
  • save_report –

  • save_path –

  • save_filename –

  • save_filetype –

Returns:

property group_id
static load(file_path)
Parameters:

file_path –

Returns:

save(file_path=None)
Parameters:

file_path –

Returns:

save_individual_dataset_reports(save_path=None, save_filetype=None)

Saves individual reports for each of the specified datasets. Individual filenames will be generated using the filename name of the dataset from which the analysis is derived.

This will result in the same analysis that can be done by creating a CaGraph object using a single dataset.

Parameters:
  • save_path – str

  • save_filetype – str (‘csv’, ‘HDF5’, ‘xlsx’)

Returns:

property threshold
class cagraph.CaGraphBehavior(data, behavior_data, behavior_dict, construction_method='stacked', node_labels=None, node_metadata=None, dataset_id=None, threshold=None)

Class for running behavior-sampled analyses on a single dataset.

This class is best suited for analyses where the time spent in each behavior is well-balanced, however, methods will be added to accommodate datasets with unbalanced behavior.

property behavior_identifiers
property data
property data_id
property dt
get_cagraph(condition_label)
Parameters:

condition_label –

Returns:

get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)

Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.

Parameters:
  • save_report –

  • save_path –

  • save_filename –

  • save_filetype –

Returns:

static load(file_path)
Parameters:

file_path –

Returns:

property node_labels
property num_neurons
save(file_path=None)
Parameters:

file_path –

Returns:

property threshold
class cagraph.CaGraphMatched(data_list, dataset_labels, match_map, matched_only=True, threshold=None)

Class for running analyses on datasets that have been cell-tracked over time to identify the same cells.

property data_id
property dataset_identifiers
property dt
get_cagraph(condition_label)
Parameters:

condition_label –

Returns:

get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)

Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.

Parameters:
  • save_report –

  • save_path –

  • save_filename –

  • save_filetype –

Returns:

static load(file_path)
Parameters:

file_path –

Returns:

property node_labels
property num_neurons
save(file_path=None)
Parameters:

file_path –

Returns:

property threshold
class cagraph.CaGraphTimeSamples(data, time_samples=None, condition_labels=None, node_labels=None, node_metadata=None, dataset_id=None, threshold=None)

Class for running time-sample analyses on a single dataset.

property condition_identifiers
property data
property data_id
property dt
get_cagraph(condition_label)
Parameters:

condition_label –

Returns:

get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)

Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.

Parameters:
  • save_report –

  • save_path –

  • save_filename –

  • save_filetype –

Returns:

static load(file_path)
Parameters:

file_path –

Returns:

property node_labels
property num_neurons
save(file_path=None)
Parameters:

file_path –

Returns:

property threshold

visualization

Authors

Veronica Porubsky

License

The package is released under the MIT License.