cagraph
Python library for generating graphs from calcium imaging data of neural activity.
Installation
Install with pip
pip install cagraph
preprocessing
cagraph
- class cagraph.CaGraph(data, node_labels=None, node_metadata=None, dataset_id=None, threshold=None)
Author: Veronica Porubsky Author Github: https://github.com/vporubsky Author ORCID: https://orcid.org/0000-0001-7216-3368
Class: CaGraph(data_file, node_labels=None, node_metadata=None, dataset_id=None, threshold=None)
This class provides functionality to easily visualize time-series data of neuronal activity and to compute correlation metrics of neuronal networks, and generate graph objects which can be analyzed using graph theory. There are several graph theoretical metrics for further analysis of neuronal network connectivity patterns.
Attributes
- datastr or numpy.ndarray
A string pointing to the file to be used for data analysis, or a numpy.ndarray containing data loaded into memory. The first (idx 0) row must contain timepoints, the subsequent rows each represent a single neuron timeseries of calcium fluorescence data sampled at the timepoints specified in the first row.
- node_labels: list
A list of identifiers for each row of calcium imaging data (each neuron) in the data_file passed to CaGraph.
- node_metadata: dict
Contains metadata which is associated with neurons in the network. Each key in the dictionary will be added as an attribute to the CaGraph object, and the associated value will be . Each value
- dataset_id: str
A unique identifier can be added to the CaGraph object.
- threshold: float
Sets a threshold to be used for thresholded graph.
- class GraphTheory(neuron_dynamics, time, pearsons_correlation_matrix, graph, num_neurons, labels)
- compare_graphs(graph1, graph2)
- Returns:
- get_betweenness_centrality(graph=None, return_type='list')
Returns the betweenness centrality scores for all nodes.
- Parameters:
graph –
return_type –
- Returns:
- get_clustering_coefficient(graph=None, return_type='list')
Returns a list of clustering coefficient values for each node.
- Parameters:
graph – networkx.Graph object
return_type – str
- get_communities(graph=None, return_type='list')
Returns a list of communities, composed of a group of nodes.
- Parameters:
return_type – list
graph – networkx.Graph object
- Returns:
node_groups: list
- get_connected_components(graph=None) list
Returns connected components with more than one node.
- Parameters:
graph – networkx.Graph object
- Returns:
list
Computes the number of connections each neuron has, divided by the nuber of cells in the field of view. This method is described in Jimenez et al. 2020: https://www.nature.com/articles/s41467-020-17270-w#Sec8
- Parameters:
graph – networkx.Graph object
return_type – str
- get_degree(graph=None, return_type='list')
Returns iterator object of (node, degree) pairs.
- Parameters:
graph – networkx.Graph object
return_type – str
- get_density(graph=None)
Returns the ratio of edges present in the graph out of the total possible edges.
- Parameters:
graph – networkx.Graph object
- Returns:
float
- get_eigenvector_centrality(graph=None, return_type='list')
Compute the eigenvector centrality of all graph nodes, the measure of influence each node has on the graph.
- Parameters:
graph – networkx.Graph object
return_type – str
- get_hubs(graph=None, return_type='list')
Computes hub nodes using the normalized betweenness centrality scores. This method sets an outlier threshold using the inter-quartile range of the betweenness centrality score distribution and returns nodes with with scores above the outlier threshold.
- Parameters:
graph –
return_type –
- Returns:
- get_largest_connected_component(graph=None) Graph
Returns a subgraph containing the largest connected component.
- Parameters:
graph – networkx.Graph object
- Returns:
networkx.Graph object
- get_path_length()
Returns the characteristic path length.
- Returns:
- get_smallworld_largest_subnetwork(graph=None) float
- Parameters:
graph – networkx.Graph object
- Returns:
float
- property graph
- property neuron_dynamics
- property node_labels
- property num_neurons
- property pearsons_correlation_matrix
- property time
- class Plotting(neuron_dynamics, time, pearsons_correlation_matrix, graph, num_neurons)
- get_single_neuron_timecourse(neuron_trace_number) ndarray
Return time vector stacked on the recorded calcium fluorescence for the neuron of interest.
- Parameters:
neuron_trace_number – int
- Returns:
numpy.ndarray
- property graph
- property neuron_dynamics
- property num_neurons
- property pearsons_correlation_matrix
- plot_all_neurons_timecourse(title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
- Parameters:
save_format –
title –
y_label –
x_label –
show_plot –
save_plot –
save_path –
dpi –
- plot_correlation_heatmap(correlation_matrix=None, title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
Plots a heatmap of the correlation matrix.
- Parameters:
save_format –
correlation_matrix –
title –
y_label –
x_label –
show_plot –
save_plot –
save_path –
dpi –
- Returns:
- plot_multi_neuron_timecourse(neuron_trace_labels, palette=None, title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
Plots multiple individual calcium fluorescence traces, stacked vertically.
- Parameters:
save_format –
neuron_trace_labels – list
palette – list
title –
y_label –
x_label –
show_plot –
save_plot –
save_path –
dpi –
- Returns:
- plot_single_neuron_timecourse(neuron_trace_number, title=None, y_label=None, x_label=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
- Parameters:
save_format –
neuron_trace_number – int
title –
y_label –
x_label –
show_plot –
save_plot –
save_path –
dpi –
- Returns:
- property time
- property betweenness_centrality
- property clustering_coefficient
- property communities
- property data
- property dataset_id
- property degree
- draw_graph(graph=None, position=None, node_size=25, node_color='b', alpha=0.5)
Draws a simple graph.
- Parameters:
graph – networkx.Graph object
position – dict
node_size – int
node_color – str
alpha – float
- Returns:
- property dt
- get_adjacency_matrix(threshold=None) ndarray
Returns the adjacency matrix of a graph where edges exist when greater than the provided threshold.
Uses the Pearson’s correlation matrix.
- Returns:
numpy.ndarray
- get_erdos_renyi_graph(graph=None) Graph
Generates an Erdos-Renyi random graph using a graph edge density metric computed from the graph to be randomized.
- Parameters:
graph –
- Returns:
networkx.Graph object
- get_graph(threshold=None, weighted=False) Graph
Automatically generate graph object from numpy adjacency matrix.
- Parameters:
threshold –
weighted – bool
- Returns:
networkx.Graph object
- get_laplacian_matrix(graph=None) ndarray
Returns the Laplacian matrix of the specified graph.
- Parameters:
graph – networkx.Graph object
- Returns:
- get_pearsons_correlation_matrix(data_matrix=None) ndarray
Returns the Pearson’s correlation for all neuron pairs.
A loaded numpy.ndarray dataset can be passed to the method for analysis, otherwise the dataset passed to the CaGraph object constructor will be used.
- Parameters:
data_matrix – numpy.ndarray
- Returns:
- get_random_graph(graph=None) Graph
Generates a random graph. The nx.algorithms.smallworld.random_reference is adapted from the Maslov and Sneppen (2002) algorithm. It randomizes the existing graph.
- Returns:
networkx.Graph object
- get_report(parsing_nodes=None, parse_by_attribute=None, parsing_operation=None, parsing_value=None, save_report=False, save_path=None, save_filename=None, save_filetype=None)
- Parameters:
save_filetype –
save_filename –
save_path –
save_report –
parsing_nodes –
parse_by_attribute – str
parsing_operation – str
parsing_value – float
- Returns:
dict
- get_weight_matrix() ndarray
Returns a weighted connectivity matrix with zero along the diagonal. No threshold is applied.
- Returns:
numpy.ndarray
- property graph
- property hubs
- static load(file_path)
- Parameters:
file_path –
- Returns:
- property neuron_dynamics
- property node_labels
- property num_neurons
- property pearsons_correlation_matrix
- reset()
Resets the CaGraph object graph attribute to the original state at the time the object was created.
- save(file_path=None)
- Parameters:
file_path –
- Returns:
- sensitivity_analysis(data, threshold=None, show_plot=True, save_plot=False, save_path=None, dpi=300, save_format='png')
Generates a series of graphs around the recommended or user-specified threshold and shows the number of edits required to transform the original graph to the series of graphs.
If many edits are required, the graphs are dissimilar.
- Parameters:
save_format –
dpi –
save_path –
save_plot –
data –
threshold –
show_plot –
- Returns:
- property threshold
- property time
- class cagraph.CaGraphBatch(data_path, group_id=None, threshold=None, threshold_averaged=False)
Class for running batched analyses.
Only directories can be passed to CaGraphBatch. Node metadata cannot be added to CaGraph objects in the batched analysis. Future versions will include the node_metadata attribute.
- property dataset_identifiers
- get_cagraph(condition_label) CaGraph
Return a CaGraph object for the specified dataset condition_label
- Parameters:
condition_label – str
- Returns:
CaGraph
- get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)
Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.
- Parameters:
save_report –
save_path –
save_filename –
save_filetype –
- Returns:
- property group_id
- static load(file_path)
- Parameters:
file_path –
- Returns:
- save(file_path=None)
- Parameters:
file_path –
- Returns:
- save_individual_dataset_reports(save_path=None, save_filetype=None)
Saves individual reports for each of the specified datasets. Individual filenames will be generated using the filename name of the dataset from which the analysis is derived.
This will result in the same analysis that can be done by creating a CaGraph object using a single dataset.
- Parameters:
save_path – str
save_filetype – str (‘csv’, ‘HDF5’, ‘xlsx’)
- Returns:
- property threshold
- class cagraph.CaGraphBatchTimeSamples(data_path, group_id=None, time_samples=None, condition_labels=None, threshold=None, threshold_averaged=False)
Class for running batched analyses on datasets that have distinct time periods to separate into samples.
Only directories can be passed to CaGraphBatchTimeSamples. Node metadata cannot be added to CaGraph objects in the batched analysis. Future versions will include the node_metadata attribute.
- property dataset_identifiers
- get_cagraph(condition_label) CaGraph
Return a CaGraph object for the specified dataset condition_label
- Parameters:
condition_label – str
- Returns:
CaGraph
- get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)
Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.
- Parameters:
save_report –
save_path –
save_filename –
save_filetype –
- Returns:
- property group_id
- static load(file_path)
- Parameters:
file_path –
- Returns:
- save(file_path=None)
- Parameters:
file_path –
- Returns:
- save_individual_dataset_reports(save_path=None, save_filetype=None)
Saves individual reports for each of the specified datasets. Individual filenames will be generated using the filename name of the dataset from which the analysis is derived.
This will result in the same analysis that can be done by creating a CaGraph object using a single dataset.
- Parameters:
save_path – str
save_filetype – str (‘csv’, ‘HDF5’, ‘xlsx’)
- Returns:
- property threshold
- class cagraph.CaGraphBehavior(data, behavior_data, behavior_dict, construction_method='stacked', node_labels=None, node_metadata=None, dataset_id=None, threshold=None)
Class for running behavior-sampled analyses on a single dataset.
This class is best suited for analyses where the time spent in each behavior is well-balanced, however, methods will be added to accommodate datasets with unbalanced behavior.
- property behavior_identifiers
- property data
- property data_id
- property dt
- get_cagraph(condition_label)
- Parameters:
condition_label –
- Returns:
- get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)
Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.
- Parameters:
save_report –
save_path –
save_filename –
save_filetype –
- Returns:
- static load(file_path)
- Parameters:
file_path –
- Returns:
- property node_labels
- property num_neurons
- save(file_path=None)
- Parameters:
file_path –
- Returns:
- property threshold
- class cagraph.CaGraphMatched(data_list, dataset_labels, match_map, matched_only=True, threshold=None)
Class for running analyses on datasets that have been cell-tracked over time to identify the same cells.
- property data_id
- property dataset_identifiers
- property dt
- get_cagraph(condition_label)
- Parameters:
condition_label –
- Returns:
- get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)
Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.
- Parameters:
save_report –
save_path –
save_filename –
save_filetype –
- Returns:
- static load(file_path)
- Parameters:
file_path –
- Returns:
- property node_labels
- property num_neurons
- save(file_path=None)
- Parameters:
file_path –
- Returns:
- property threshold
- class cagraph.CaGraphTimeSamples(data, time_samples=None, condition_labels=None, node_labels=None, node_metadata=None, dataset_id=None, threshold=None)
Class for running time-sample analyses on a single dataset.
- property condition_identifiers
- property data
- property data_id
- property dt
- get_cagraph(condition_label)
- Parameters:
condition_label –
- Returns:
- get_full_report(save_report=False, save_path=None, save_filename=None, save_filetype=None)
Generates an organized report of all data in the batched sample. It will report on the base analyses included in the CaGraph object get_report() method, and output a single pandas DataFrame or file which includes these analyses for all datasets in a tabular structure.
- Parameters:
save_report –
save_path –
save_filename –
save_filetype –
- Returns:
- static load(file_path)
- Parameters:
file_path –
- Returns:
- property node_labels
- property num_neurons
- save(file_path=None)
- Parameters:
file_path –
- Returns:
- property threshold
visualization
License
The package is released under the MIT License.