統計鹼基數目、GC含量、read數、最長的read、最短的read及平均read長度

本文轉載自查看原文 2017-01-19 13:14 4584 數據處理小代碼

# 用於fasta格式文件的鹼基數目和GC含量的統計

grep -v '>' input.fa| perl -ne '{$count_A=$count_A+($_=~tr/A//);$count_T=$count_T+($_=~tr/T//);$count_G=$count_G+($_=~tr/G//);$count_C=$count_C+($_=~tr/C//);$count_N=$count_N+($_=~tr/N//)};END{print qq{total count is },$count_A+$count_T+$count_G+$count_C+$count_N, qq{\nGC%:},($count_G+$count_C)/($count_A+$count_T+$count_G+$count_C+$cont_N),qq{\n} }'

# 用於fastq格式文件的read數、鹼基數、最長的read、最短的read及平均read長度

perl -ne 'BEGIN{$min=1e10;$max=0;}next if ($.%4);chomp;$read_count++;$cur_length=length($_);$total_length+=$cur_length;$min=$min>$cur_length?$cur_length:$min;$max=$max<$cur_length?$cur_length:$max;END{print qq{Totally $read_count reads\nTotally $total_length bases\nMAX length is $max bp\nMIN length is $min bp \nMean length is },$total_length/$read_count,qq{ bp\n}}' input.fq

# 用於fasta格式文件的read數、鹼基數、最長的read、最短的read及平均read長度

perl -ne 'BEGIN{$min=1e10;$max=0;}next if ($.%2);chomp;$read_count++;$cur_length=length($_);$total_length+=$cur_length;$min=$min>$cur_length?$cur_length:$min;$max=$max<$cur_length?$cur_length:$max;END{print qq{Totally $read_count reads\nTotally $total_length bases\nMAX length is $max bp\nMIN length is $min bp \nMean length is },$total_length/$read_count,qq{ bp\n}}' input.fa

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 談談MySQL的基數統計最長最短單詞 mysql 統計查詢出來的數目 codevs 1862 最長公共子序列（求最長公共子序列長度並統計最長公共子序列的個數） R 統計True/ False的數目平均查找長度平均路徑長度與直徑最短路 || 最長路 || 次短路（算法）無向圖最短路徑的數目 Python 獲取文件中最長行的長度和最長行