I need a web scraper written for the .xlsx file in the following directory:
[login to view URL]
The latest .xlsx file within that directory will need to be downloaded.
The name of the file is subject to change daily and will need to be identified by the latest .xlsx extension.
All information needed is available on the main page. The number of rows will vary.
The output should be a pipe (|) delimited file with the following column mappings:
origin_city --> data located in the "Origin City*" column (column B)
origin_state --> data located in the "ST" column located after the origin city column (column C)
ship_date --> data located in the "P/U Date" column (column A), change to the YYYY-MM-DD format
destination_city --> data located in the "Dest. City" column (column D)
destination_state --> data located in the "ST" column located after the dest. city column (column E)
receive_date --> leave blank
trailer_type --> data located in the "Eq. Type" column (column F)
load_size --> data located in the "Full/LTL' column (column G)
weight --> data located in the "Weight" column, add three zeros to the end of the
data if there are only 2 numbers (ie. 48 = 48000) (column O)
length --> data located in the "Length" column (column N)
width --> leave blank
height --> leave blank
trip_miles --> data located in the "Miles" column (column H)
pay_rate --> data located in the "Rate" column (column M)
contact_phone --> leave blank
contact_name --> add data from "Dispatcher" and "Ext." columns (columns J & K)
tarp_required --> leave blank
comment --> data located in the "Special Req." column (column L)
load_number --> leave blank
commodity --> data located in the "Commodity" column (column I)
The first line of the output should contain all of the column headers.
Any field that contains no data should be left blank.
Please do not use words like "null" or "blank" in blank columns.
Below is a sample output of the first 5 columns using sample data:
The deliverable will be a Perl .pl file that must run on
Ubuntu Linux and must use Modern::Perl. The Perl .pl file
should be called '[login to view URL]' and the output file should be
called '[login to view URL]'
It will be scheduled in cron to run unattended every 15 minutes.
Please specific what language/OS/modules you plan to use.
Also, please include the word "raccoon" in your bid so I know that
you read this description.